Quantum Workforce Intelligence from qubitsok.com with Piotr Lewandowski
E108

Quantum Workforce Intelligence from qubitsok.com with Piotr Lewandowski

Summary

Who is actually doing quantum computing research, where are they, and what does the industry really need — not according to press releases, but according to the data? Piotr Lewandowski built qubitsok.com to answer exactly those questions, constructing what may be the most granular open-access dataset of quantum jobs, researchers, and publications in existence. This conversation explores what that data reveals about the real size, concentration, and trajectory of the quantum workforce — and why the answers are more unsettling than the funding announcements suggest.

Sebastian Hassinger • 00:04
This is the New Quantum Era, and I'm your host, Sebastian Hassinger.

Sebastian Hassinger • 00:30
Soon after I joined IBM Quantum in 2018, I developed a fascination boarding on obsession with the archive, where exploration and progress and our collective mastery of quantum information technologies could be seen unfolding in real time. The archive site, found at arxiv. org, has, since 1991, been a place to post and share preprint research papers. Initially hosted in a shared electronic mailbox at Los Alamos National Lab, the archive evolved and migrated into the website hosted by Cornell University since 2011. The category covering quantum information is called quant-ph, and over 1500 papers are posted to it per month What I found tantalizing about the archive was the idea that not only piping hot research from the absolute bleeding edge of the field could be found in the papers, but that the corpus as a whole represents an unbelievably rich data source. Offering insight into the evolution of the field and the people who are driving it. It's the people in particular that I find fascinating And thinking about those people is something the emerging industry needs to do a lot of so we can build a better idea of what the workforce actually looks like. Not the workforce that gets described in funding announcements or government roadmaps, but the real one, the people publishing papers, taking jobs, moving between countries. contributing to open source libraries, and even deciding whether to stay in quantum or move on.
My guest today, Peter Lewandowski, has been building the infrastructure to answer that question. He's the founder of qubitsok.com, which started as a tool for researchers, evolved into a quantum job board, and is growing into something considerably more ambitious. a continuously updated intelligence platform that ingests the entire quant pH archive, tracks thousands of open quantum rules across hundreds of companies, parses patents and grants and open source repositories and runs all of it through a custom ontology of 500 plus tags that he built from scratch. He does this as a side gig from his home in Poland, which makes his work that much more impressive, and also raises interesting questions of what it means to be an outsider analyst in a field that's very good at talking to itself.
I've been a subscriber to his daily paper digest for a while now, and it's one of the few things in my inbox that I actually read religiously. The tagging is really useful. The summaries are better than what you get from a quick abstract scan. And occasionally the data surfaces something that makes you stop and reconsider what you thought you knew about the field. When I joined AWS and was trying to get some hands-on knowledge of their cloud services, I built a very rudimentary scraper just to try to track keyword trends in the research output. What Peter has built makes that look like a napkin sketch. He's not just ingesting the entire quant-ph. He's tracking author affiliations and normalizing them to country and city level. He's following citation networks and co-authorship graphs over time. And he's starting to use all of that to do something quite novel, match quantum job optings to specific researchers based on claim-level evidence from their published work. not just keyword overlap, but actual verification of whether someone has hands-on experience with, say, ion shuttling or FPGA decoder engineering. In this conversation, we get into how the ontology works and why he built it the way he did.
We talk about what the data shows about researcher mobility, which countries are gaining quantum talent, and which are losing it. And some of the answers are not what you'd expect. We talk about the rising share of industry authorship in the research literature and what that structural shift might mean for the science. And we talk about a new tool he's building, which he's calling "qubie," that he thinks could fundamentally change how quantum hiring works. There's also attention running through all of this that I want to explore. A platform that gets good enough at identifying and ranking quantum talent starts to have its own gravitational pull on the field. When you can measure something this precisely, you inevitably start to shape it. I'm curious whether Peter has thought about that and what he makes of it. Here's my conversation with Peter Lewandowski.

Sebastian Hassinger • 05:07
Peter, hello. Thank you so much for joining me. we've talked many times and I've been thinking about getting you on the podcast for a while now, so I'm really happy to be having this conversation today. You're the force behind qubitsok. com, which I'm fairly obsessed with. for me

But all the way back to IBM, I was really interested in the archive and quant-ph specifically on the arxiv because There's such a high volume of research output on a daily basis, and the field is still very oriented towards open science and sharing what people learn.

and the affiliations and the author names and their age score and all the rest of the metadata around the contents of the research papers. is such interesting source of information from my perspective. When I joined AWS, to try to learn some of the cloud services, I built a very rudimentary scraper or the archive I pulled from their various sources, you know, public sources and try to do just some static searches of keywords, either, you know quantum hardware vendors or particular algorithms. And I found that really fascinating. So when I found what you were doing at qubitsok. com, I was kind of like I had one of those, you know, brain exploding kind of moments because what you're doing is so incredibly rich and there's so much interesting insight in the daily newsletter that I get from your site. So start with

What was your original inspiration for building qubitsok. com? Where did you start with it and how did you go about building what you've done so far?

Piotr Lewandowski • 06:52
Yeah, thank you. But thank you for having me. and indeed we had a couple of great conversations beforehand. and like thank you for all the support for qubitsok from the not very beginning of the qubitsok but from the very beginning of like our conversations so basically originally qubitsok started as a job board

I was in a place when I was looking for a job myself. As a software engineer and I like dabbled on the theory part of quantum computing beforehand, and I wanted to

like have a proper categorized jobs for quantum computing. That's one thing. And another thing is it was a moment, one of those AI moments when

Well, I guess this happens now every few weeks. but AI was really started getting popular and I really wanted to get into AI assistant assisted coding. So with those two things, I started building a couple of scrapers to understand the lay of the land when it comes to the jobs. And then like people started to using it Yeah, which was very like fun and I got a lot of positive feedback. And the next thing was, hey, maybe I can like

just to give back to the quantum computing community. Build a

newsletter or a view for archived paper so they are tagged a little bit better because right now well right now originally what we

happens to have is just quant-ph. That's all and every single quantum computing paper goes through that single hose. I decided to Start on a continuous basis, parsing all of the archived papers and started tagging it. with various types of research from hardware to software, the modalities of qubits, type of algorithms so people can select

Whatever is of specific interest to them, and they will they could get a originally it was an abstract of the paper, but right now it's even a shorter summary for each of the paper and all the tax that were that were assigned. So with that it started

Growing slowly but surely, and like I started just with grabbing the papers. What I started doing from the very beginning is Reading full PDFs, not only abstracts, so that the tagging is usually quite good.

And the summaries of the papers, very short summaries, are also usually quite good. They are still not perfect, but I'm pretty sure in a couple of

months they will be absolutely perfect. They are still like not not to be like too fool of myself, but they are still or already better than anything else and they are like let's say 95% correct.

And there are a couple of things that hallucination do do occur, but I am constantly like improving and refining the system

So that's how it all started. and then I started gathering more and more Oh those papers. Yeah.

Sebastian Hassinger • 11:11
Right. But let me just stop you for a second because I'm I'm really curious about a few things. So you do you call semantic tagging. I think you know, that kind of leads to a kind of ontology that you're building, right? I mean you're building an ontology that underpins

the quantum research papers and sort of describes the network of or the system of sense making or knowledge making that is embedded in across all those research papers. Was that process did you initially

Think of that as being a way to help people search better or help the AI produce better results in its parsing and its sort of summarizing?

Piotr Lewandowski • 11:58
So when I started building ontology of well quantum computing, right now there are like 500 tags and they are all

forming a tree-like structure so we have a parent-child relationship between between the tags originally I wanted to steer AI into

using the most specific tag. And yeah, it's actually it was quite successful from the very beginning, even with the older older generation of models. and w because with this I didn't want for AI to bother its own context with something like

If this paper relates to quantum hardware, it's obviously it is valuable, but the most valuable thing is like when we are talking about the hardware what type of modality are we proposing something Are we doing overview of the current technology? Are we doing experimentation when it based on particular qubit modality or specific hardware? So

This was the reason why I wanted to use this ontology and this is why it is in a tree-like structure.

Sebastian Hassinger • 13:36
Right, right. And then so you were about to say you started ingesting more and more papers. At this point, you've gone back, I think, to the beginning of quant-ph, right?

Piotr Lewandowski • 13:47
So yes, I did. there are I do not recall precise number.

So but like I've parsed whole quant-ph and couple of already adjacent fields.

But those adjacent fields like for example condensed matter, it is happening in a two-way step. So first what I do, and this is based only on the abstract I check whether this particular paper relates to the quantum computing in any way, shape or form and then only if it does then I will extract all the tags based on the full PDF. So this means that from the beginning of archive I've parsed

I cannot say all of them, but let's say ninety-nine point nine percent of the papers that have

Sebastian Hassinger • 14:56
some relation well or not some but direct relationship with yeah that's a huge core and you've even branched out to other sites as well I think right

Piotr Lewandowski • 15:08
so yes so this depends so what I did is that right now

When it comes to full papers, I'm using only archive for that, but I started using more and more sources to gather more and more data. So right now the database yeah yeah I will stop now.

Sebastian Hassinger • 15:39
Yeah. So you're doing

enhancement with additional metadata sources then that refer to those same papers. Things like like SCIRATE or Google Scholar or something like that you're pulling additional or sort of metadata about the authors or the re the references, citations, that sort of thing. Is that right?

Piotr Lewandowski • 16:02
So yes, the core part is citations. So I do extract on a continuous daily basis all these citations for the papers. which is a like quite intensive process. But right now from so for each of the paper I also do extract the author, their affiliation, if they if they provide it within the paper. It's not like I am dispatching like I don't know billions well well not billions but millions of queries to find every person every day. But if someone discloses it within the paper, then I will get this affiliation. I will also extract from the I will normalize the affiliation, check in which country it is, so we can have like a country level data or city level data to understand where where research is happening when it what when we are talking about geography. and citations and co-autor co-authorships can also give a absolutely unique dynamic like overtime understanding of how the field was changing and who is doing what right now who is Well and this please take it with a grain of salt, who is like a rising star meaning publishing papers that are sure getting cited.

Sebastian Hassinger • 17:42
Yeah, I mean by great assault you mean it's you have created criteria on a somewhat arbitrary basis to identify whatever a velocity of papers that are coming out and numbers of citations that are are piling up and sort of made a dividing line and saying anybody above this level is a rising star, right?

Piotr Lewandowski • 18:04
Yeah, yeah, yeah. Exactly. I just don't want to say that like like

I don't feel strongly enough to say that someone is. A rising Starbucks is based on their same citations.

Sebastian Hassinger • 18:17
I under I understand. and so that's just as it as you've described it so far, it's a very powerful

search and research tool for the community, right? I mean it's not just simple plain text searching, but it's semantic tag enabled searching. there's a lot of metadata added. There's a lot of organizational structure that's that's derived from the content of the papers. Is that sort of the was that kind of the first I guess feature set that you were aiming for was sort of a researcher's dream tool for how to search the active body of research papers

Piotr Lewandowski • 18:59
Yes, that was the idea for like the first first iteration, first version Another thing that is also freely available to the community right now, if they will go to qubitsok.com slash collaborate. you can search for people by their like expertise based on the ontpology affiliation so let's say you are writing a paper on a topic X

Yeah and the and you would like to find someone who maybe has similar experience or you need you would like to talk to someone with adjacent experience to double check something or brainstorm whatever then you can use this tool and find those people. Even you can specify affiliations so you can find people from universities you your or universities or companies that you will find interesting to work with.

Sebastian Hassinger • 20:03
That's really cool. And so

That's kind of in the context of tools for researchers within the field. But obviously when you do that much data analysis with that large a corpus, you're going to start uncovering things that are sort of about the field itself rather than about the individual threads of study within the field. So you've you've really started to uncover demographics, geographical trends, trends over longer periods of time, organizational affiliation trends. What's the sort of shape of the under your understanding of the field of quantum research itself has started to emerge at that metadata level.

Piotr Lewandowski • 20:50
That's a great and tough question.

That's my specialty, Peter. Yeah.

Yes, so the field is getting bigger and the velocity of paper getting published is increasing even before let's say AI spring it is

quite clear that the field is booming. I it will be interesting for me to see how it will go over

next couple of orders to understand whether

Researchers are still joining the field on a similar like velocity as they did before So that's that's one thing. another thing that I can like see that in last twenty-four months the

Biggest country that gained quantum computing talent are Germany Which I think is not some not something someone could expect. And obviously this is mobile mob mobility, what does it mean? It means that someone moved like used to publish under institution in one country and then started publishing for a prolonged period of time in institution into another country So that's that's the that Germany is a big winner. China, obviously well obviously or not, but they are gathering more and more talent. the biggest countries that are getting brain drained are interestingly

Or not. United States.

It seems that in the last two years US lost a little bit of them researchers. the Australia that's interesting for me also, but also Poland is in top five So it means that like for Poland probably researchers were like groomed and raised in Poland and then taught and then

And I will postulate because I'm I'm just as a disclaimer, I'm Polish and based in Poland. and because the ecosystem itself it's not just as mature here when it comes to especially like Industry-wise. So probably the talent is living to a place is more friendly to that. Another thing that is clearly to be seen in the data is the share of the industry publishing meaning that

Like 10 years ago, I think let me can I double check it quickly?

Sebastian Hassinger • 24:10
industry authorship on quant pH papers was 3. 4% in 2005, so 20 years ago And is now fourteen percent in twenty twenty six.

Piotr Lewandowski • 24:26
Yeah, yeah.

Sebastian Hassinger • 24:30
Makes sense. Yeah, it does. But it's a r it's a very good indication of the health of the industry. And but before we move on, I just wanna so when you talk about migration, either brain drain or in you know, growth of the quantum community, are you able to sort of I mean it's very hard to establish the size of the of the population of the sort of you whatever, qu industry qualified quantum researchers, but you kind of can make a proxy from from the authorship of the data.

Piotr Lewandowski • 25:02
That's physical that's absolutely

feasible we can but basically using the affiliation as the proxy for the country Meaning that we can normalize and then re represent brain drain as a percentage the of the whole population.

Unidentified speaker • 25:24
I do not have this number right now.

Piotr Lewandowski • 25:27
But the numbers or sorry, the countries I've mentioned before, those were positioned by the absolute numbers.

Sebastian Hassinger • 25:36
That's so interesting. And the industry trend line I think is very interesting. I mean it what's interesting to me is that

There's sort of two countervailing forces. On the one hand, there's industry, you know, private sector investment and growth. hiring people out of the academic community who then start publishing with the affiliation of a company. But then of course there's also the IP concerns driving research sort of more, you know, in a proprietary direction and less disclosure. Do you Do you see, is there any ability to see like whatever, like somebody moves from a university to a company and then their volume drops off?

Piotr Lewandowski • 26:24
We that happens.

We can observe that usually when people join industry

the number of papers they publish tend to drop and but

Also, like there are a lot of caveats to that, because it depends on the on the company, the role, and so on and so forth. But

On the other hand, what happens with the within the let's say private sector. For example, we have ability to see through their open source contributions. because what qubits okay tracks also is the number of basically i parse most of the open source libraries related to quantum computing to understand also who is contributing the most and where.

Sebastian Hassinger • 27:29
That's incredible. And I mean what you just described, all of that is in the sort of in the open. It's in your your free to use qub tsok.com

But I mean what you're describing feels like the tip of an iceberg in terms of value that you could create with this data and it and it starts to sound like an analyst product, right? I mean people like Quantum Insider or GQI or Bob Sutor or others are providing analysis, professional analysis services. Is that some something that you're considering doing or you're working on d that's a very good question.

Piotr Lewandowski • 28:07
obviously there is maybe even to a little make it even more interesting. Not only do I parse all of the archives

data related to quantum computing. Not only do I get open source contribution, but I also ingest all the patents. I do ingest all of the

grants that I can put my hands on. So like which institution is getting funded. and anything else that is cool to mention? Those are probably the coolest.

And also a couple of company-related events like like this financial stuff around anything, equity and stuff like that. So

I think qubitsok has absolutely unique person-first view of quantum computing industry integrated across multiple sources. In total there are 40 different sources because yeah like even grants there are I think 30 different ones. So that's a vet. And there are

I see like two main lanes but the qubits okay can go on the business side because what I want to

always provide for free are tools for researchers and for people looking for jobs. I never want to get money from them. I don't want to like make it somehow

So yeah, so there are two main lines of the product that at least two that one can build. One is analytics or

Focused and another is like talent mature focus. And as you know, because we discussed it a couple of times, it's

Not an easy thing for me to choose which way to go. But right now I decided to pursue at least temporarily, mostly the talent matching part. at least publicly which means that

Because if I have all of research done in quantum computing and you are looking for a quantum talent

This gives you a unique opportunity to try to get to those people and to understand what they are researching on. And moreover, because of the citations, probably you have already a scientist in your company. So what you can basically do using qubits okay data is to try to get a world intro. Maybe someone you call for the paper couple like couple year years before knows this person and you forget you forgot that you do it that they know them because Maybe they d like it's happened over time, right? All of those networks I dynamic Yeah.

Sebastian Hassinger • 31:36
It's almost like the it's like the whatever, informal network or the weak ties. That's the that's the social science. Science's way of describing those sort of passing social relationships that have s you know surprisingly high value. in frankly in business settings. So in a way what you're doing is your data is surfacing those weak ties.

Piotr Lewandowski • 32:02
Yes. And Obviously, initially most of the startups will rely on the direct network o of the core scientist.

But as they grow and they need to after, before or during getting their funding, they need to understand the talent pool they are working with So we can not only present the absolute number of people under a specific criteria but we can show who you can reach real realistically and that's pretty awesome but

Right now, what I am building is because it's an AI era and I don't want to say that it's an agent, but it is a tool that uses AI or models, and basically But I'm calling it qubie and I guess this will be a first time with this menu anywhere. But what qubie does

is let's say that you have a role as there are many roles on qubitsok.com slash job the

We can take this role, understand where you are looking for people, and what type of people you are looking for, but what type I mean like what type of research they needed to perform beforehand

And what you can do initially is basically calculate tag overlap using one of the methods that allow to you to compare the tree-like structure but that's that's one thing. But then what you can do is to pres to dispatch a couple of agents

and extract claims from your job. So let's say that you are looking for someone that has Experience with I don't know iron shuttling or yeah let's say with iron traps So and you want something someone with hands on with iron traps with someone actually like either building or using ex using iron traps. So what we can do

Is to we can dispatch those sub-agents to read through all of their papers, through their dissertation, through their different places, maybe they have a blog

So think about all of the unstructured data related to dispersion. And we can extract following information about the claim. Like they did work with iron traps.

Or we do we well there are three options. We did not prove anything, so you would need to ask them. during a conversation hey have you have you had opportunity to work with iron traps and if so

And if so, what exactly type of work was that? Or which would be very hard to find, but in theory it is possible, contrafactual, maybe person I don't know, had written that. I will did not work with iron traps and I will never work with iron traps in my life. So then you have a like a negative statement about this So with that you can compose a profile of a person.

Something that nobody has time for that.

There is two issues. Either if you are running a company and looking for such talent, you don't have time to go to I don't know, gather 50 people first

Like because what you are looking for a talent to fill in a specific hard to fill quantum computing frame. First, you need to build a pool of people that even make

make provisional sense when it comes to filling in this role and then it would be great to investigate each of their paper, right? To understand what they did. Maybe they have a PhD dissertation where many people will write exactly what they worked with during their PhD. So you can understand what tools they used and what what they are hands-on and what qubie does is aggregates all of this data for any anything between twenty five to fifty people for your particular role and then wow it will give you back a report saying if you want to talk to this person they did this they did they did a b c and b I could not find information about B, C or sorry, XYZ.

So please during your conversation Ask them this question to understand how they would so an interview guide in a way. Yeah, basically basically like I like to dream semi-big or big My idea is a bit outlandish and it is to ensure that sourcing in quantum computing is solved Not technical recruitment, not the recruitment itself. It is way too important to leave it to AI. But

Fun fact my first job during college was a sourcing I don't know sourcing intern or sourcing partner and

Very big and famous company starting with H and ending with S.

And

Sourcing work is for me it was a very soul crushing thing. Because what you did is you went through LinkedIn You went through profiles with for me it was pen and paper and I was like writing names and then seeing whether they matched the profile the needs then There was a call that like nobody like nobody most of the people don't want to have this conversation which is very weird about the experience. Yeah. And research has this unique like g gives you this unique lens to see people's work before you talk to them. Right. And you can

You can be really prepared and this helps on two sides. Because on the first side, you will talk to people that are actually capable of filling in the role that you need very specifically with claims proven or unproven. And then when you are talking to them, We're not wasting their time asking questions that they already answered via their published work.

Sebastian Hassinger • 39:10
Right.

Piotr Lewandowski • 39:10
So it's the win-win scenario, there is less of the soul crushing profile going work There is you give you get more information, you can quick more quickly and more efficiently connect with those people.

And on the other hand, while you are talking to them, you can be like way more prepared than anyone can like I've done it like for a couple of my roles, my roles, roles on the cubis, and the results are pretty pretty awesome. And I think this

will be a truly game-changing moment when it comes to connecting like

Sebastian Hassinger • 39:58
companies with talent. That's really great, Peter. I mean what I love about this is not only is it, you know, this from the starting point, I had this fascination of what what secret y knowledge and learnings could be hidden in that giant pile of data on the archive about the field itself, but then you're turning that around and focusing on the people who are producing all that incredible work. and using AI for the boring part, not the interesting part. Like having a conversation like this is one of my joys. And I would hate to have you know an AI podcast host being the interviewing you. I'd rather do the research with help with AI and then be well informed when I have the conversation with you. So hopefully I come up with some good questions. And that's what you're enabling recruiters and hiring managers to do. And on top of all that, you're doing that at a time when There's a lot of concern and interest in workforce development and what happens when this industry starts to get traction. and suddenly there's more open positions for very highly specialized skill sets than there are people or maybe ha have ever been people to fill them all. So it feels like a an incredibly important thing you're doing, I have to say. Thank you.

Piotr Lewandowski • 41:17
What else I can say?

The other other potential line of product is

the basically either VC or competition research or analytics Which the data is there and like if someone needs a good reaper, like drop me a message on LinkedIn.

we can we can think about something but in general I want to focus on connecting talent because honestly I just think it's way cooler and way

way more helpful for people for like community let's you are let's say you are a researcher that is doing a Great research in a particular university, over and over and over again. when you read about experience of people in academia when they are thinking about joining industry they are not sure about their skills even if they are great it's whatever

Whatever composition of human psyche is there, unfortunately many of great Absolutely intelligent people are doubting themselves, I think, because they know because of their like experience, like basically done proper effect. Because they have so much experience they understand what they don't know and they think like there are and usually you can have a colleague at the department that maybe

It's s it doesn't matter like whether they are quote unquote smarter smarter or have more popular research or whatever. You can your brain can play so many tricks on you. That's right.

you can and it's gonna work both ways. First I will do reach out to the companies when this product will be ready in a couple of weeks. or yeah, like in a small couple of weeks, so they can collect them themselves to those researchers

and they can like go and build very cool stuff in the internet

I think it's a good moment to build a lot of cool stuff in the industry.

Sebastian Hassinger • 43:53
It definitely is. So Peter, thank you so much for joining me. I think what you're doing is really great. I look forward to you know my my report in my email inbox every day. and I encourage everybody else to go check out qubitsok.com and When you develop a new feature or a new product, I'm gonna have you back on and we can talk about it some more. So thank you so much. Thank you. Thanks so much to Peter Lewandowski for joining me today. One of the things that

Going to stay with me from this conversation is something he said almost in passing that sourcing work is soul crushing when you're doing it by hand, going through LinkedIn profiles with pen and paper, and that the published research record is actually far richer and more. honest signal about what someone can do than anything a cold call process surfaces. That's a simple observation, but it points to something real about how mismatched our hiring infrastructure is with the actual nature of quantum expertise. The work is public. The evidence is there. We just haven't had the tools to read it systematically until now. If you want to explore what Peter has built, go to qubitsok.com. The job board, the paper digest, the researcher collaboration search are all free. Free. If you're a company looking for quantum talent, reach out to him directly on LinkedIn. He's actively building out the qubie talent matching product and is looking for early partners to work with. Full show notes for this episode, including links to either salary report.

His skills analysis, his peer-reviewed paper on the quantum job market, and a landing page for qubie are at newquantumera. com While you're there, sign up for the newsletter. It's the best way to stay current on new episodes and the occasional longer form piece we put out. If you're getting value from the new quantum era, the single most helpful thing you can do is subscribe on whatever platform you use, Apple Podcasts, Spotify.

YouTube, Amazon Music, and share an episode with someone who's trying to understand where quantum is actually going, not just where the press releases say it's going. Word of mouth is generally how this show grows. Thanks for listening. I'm Sebastian Howard.

And this has been The New Quantum Era. Theme music by OCH. See you next time.

Creators and Guests