One page of definitions and limits, so the chapters do not have to repeat them. Worth ten minutes before quoting any figure, because most of the alarming things written about this research come from misreading one word in it.
Every rating in this report answers one question, put to human annotators. Their words, not ours:
Would access to a large language model, or software built on one, reduce the time it takes to complete this task, at equivalent quality, by at least half?
Read it twice. It is a claim about how long a task takes. It is not a claim that the task stops needing a person, that it will be done by a machine, or that a job disappears. Nearly every frightening sentence written about this research comes from collapsing the first into one of the others.
"At equivalent quality" is doing work too. The measure is speed at the same standard, not a better result.
Every task falls into one of three. Charts throughout the report use the same three fills in the same order.
A chat window on its own can halve the time, at the same quality. Nothing to install, nothing to build, available now. This is the smallest of the three bands almost everywhere.
A chat window alone is not enough, but the annotators judged it easy to imagine purpose-built software that would halve it. That software may or may not exist. This is the largest band in most occupations and the softest of the three, because "easy to imagine" is a judgement, and it is where two competent raters disagree most.
No halving is possible at the same quality. The rubric is explicit that work requiring a high degree of human interaction belongs here, which is why so much of what survives across the report is a conversation with somebody.
Where a single figure is quoted, it counts reachable today in full and needs software as a half, because that software is not confirmed to exist. Wherever there is room, the report shows the three bands instead, because one blended number invites "so AI does most of my job", which is not what any of it measures.
Nobody surveyed an employer. Exposure is about the task, not about headcount, budgets, or what any organisation intends to do.
Tasks are counted equally. An occupation's list makes no distinction between something done for five minutes a month and something done for three days a week. Every percentage here is a share of listed tasks and never a share of time.
These are occupation-level ratings against a standard task list. Your job contains tasks that are not on it, and skips ones that are.
O*NET is a US taxonomy and the ratings were made against US task lists. But the thing being rated is a task, not a country. Drafting a policy document, interviewing an applicant or reconciling a payroll is the same work in Manchester as in Ohio, and the rating is about what a model can do with that work.
What does not travel is the mix. Which occupations exist in what numbers, how duties get bundled into one role, and every legal and regulatory specific. So the task-level bands should hold across advanced economies, and the occupation-level figures are approximate outside the US rather than wrong.
There is some support for it. Google's ATLAS study, published in July 2026, mapped 15 million Gemini interactions to the same O*NET task list across more than 150 countries and 140 languages. A study that finds the same task structure useful across almost the whole world is evidence that the tasks themselves are not a US artefact.
It remains a reasoned position rather than a measured one. Nobody has re-rated exposure against UK or European task lists, and until somebody does, this is an argument.
The ratings were made against the model capabilities of that year. Treat every figure as a floor. Where a task looks obviously reachable today and is rated otherwise, that gap is the age of the research showing through, not an error.
The exposure ratings here come from human annotators applying a written rubric to task descriptions. Microsoft measured something similar from the opposite direction: 200,000 anonymised Microsoft Copilot conversations, scored for how much of each occupation AI was actually being used for.
Different data, different method, different year, no shared authors. If both are measuring something real they should agree. Across 773 occupations in both studies:
That is a strong agreement for two studies with nothing in common but the occupation list. It does not make either one right, and it will not rescue a figure that rests on twelve tasks. What it does is make the shape of this report trustworthy: which kinds of work sit at the exposed end and which do not.
The two scales are not comparable and are never shown side by side. One counts tasks against a rubric; the other scores occupations from observed use. Only their agreement on ranking is meaningful, which is what a correlation measures.
They disagree on level for a reason worth knowing: every usage-based measure samples the people who use that particular product, and those populations are not the workforce. That is an active research problem, not a detail. It is also why this report is built on the rubric-based ratings and uses the usage-based one only to check them.
Tufts University published its own occupational exposure index in March 2026, built by combining three separate academic measures and then layering on what people actually do with Claude and Copilot. It covers 757 occupations, keyed the same way as ours.
Across 757 occupations in both, their index and our share of work an AI can reach at all agree at r = 0.91, and on rank order at 0.92. On the share of work that cannot be sped up, the correlation is -0.91, which is the same agreement pointing the other way.
This one is not independent, and that matters. One of the three measures they combine is the same Eloundou rating this report is built on, so roughly a third of their input is our input. What the agreement really shows is that the other two measures, from Brynjolfsson and from Felten, largely agree as well. Treat it as three methods converging rather than as two studies confirming each other.
Their work goes somewhere this report will not, converting exposure into projected job losses by city and by state. Those numbers are theirs and are not repeated here. We have written separately about the range behind them.
Every occupation in this report carries a Job Zone, which is O*NET's measure of how much preparation a job needs before somebody can do it. It combines three things: the education normally expected, the previous experience required, and the on-the-job training it takes to get up to speed.
It is not a measure of difficulty, importance or pay. A Job Zone 2 job can be physically punishing and a Job Zone 5 job can be dull. What it describes is the distance between hiring somebody and them being useful.
Usually a high school diploma or GED, though some occupations may not need one. A few days to a year of on-the-job training.
Usually vocational training, related on-the-job experience or an associate degree, plus one or two years learning alongside experienced workers.
Most require a four-year degree, but some do not. Several years of related experience or training on top.
Most require graduate school, and some a doctorate, a medical degree or a law degree. Many need more than five years of experience.
Read down those bars and the finding this whole report rests on is visible in one place. The more preparation a job needs, the more of it an AI reaches. Seventy-six per cent of the work in the least prepared jobs cannot be sped up at all. In the graduate band it is thirty.
That is the opposite of every previous wave of automation, which came for physical work and left thinking alone. It is also why "will a robot take my job" was the wrong question to inherit.
Zone 1 is not shown because no occupation in this report has enough rated tasks to clear the floor. O*NET groups zones 1 and 2 in its own guidance. O*NET's full definitions .
The report covers 19,265 tasks, which sounds like plenty, but they are spread across 923 occupations. The median occupation has 20 rated tasks, and some have as few as five.
So a figure like "8 per cent of this job cannot be sped up" can mean one task out of twelve. It reads precise and it is not. Wherever the report shows an occupation, it shows the task count beside it, and the task-square charts are drawn so that a row's length is its sample size.
Trust the large gaps, not the small ones. The difference between 76 per cent and 30 per cent is real. The difference between 19 per cent and 21 per cent is not.
There is no fixing this by finding more tasks. 19,265 is the entire O*NET taxonomy, so the sample is the population. What can be done instead is checking the answer against an independent study, which is the section above.
"An AI tool scored this role at 58" is not an objective, fair and consistently applied selection criterion. It would not survive a tribunal asking how the pool was chosen, and it would not deserve to.
The report exists to show where to invest in training and where to build something, never who to remove. If a figure from here is being used the other way, it is being used wrongly, and no wording on this page changes that. It is worth saying anyway.
| Source | What it provides | Licence |
|---|---|---|
| Eloundou, Manning, Mishkin and Rock, GPTs are GPTs: Labor market impact potential of LLMs, Science 384, 1306–1308, 2024. Data | The exposure rating for every task. Human annotators, 19,265 tasks | MIT |
| Tomlinson, Jaffe, Wang, Counts and Suri, Working with AI: Measuring the Applicability of Generative AI to Occupations, 2025. Data | An independent exposure measure used only to check ours, never mixed with it | CC BY 4.0 |
| Google, AI & Economy ATLAS v1.0, July 2026. 15 million Gemini interactions, 150+ countries, 140 languages | Cited only, for its geographic coverage. No dataset is published, so nothing from it is joined here | Report only |
| O*NET Database version 30.3, USDOL/ETA | The task statements, occupation titles, real job titles, and the education held by people doing each job | CC BY 4.0 |
There is no data in this report on what anyone is actually doing with AI. A dataset exists that appears to measure it, and an earlier draft used it. It was removed because its columns are not defined in any documentation we could find, and the figures its publisher has quoted cannot be reproduced from the file under any aggregation we tried. A number whose meaning cannot be stated does not belong here. If that changes, this page will say so.
Everything in the report is therefore about what published research judged to be possible in 2023. None of it is evidence about practice.
This report includes information from the O*NET Database by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA). Used under the CC BY 4.0 license. O*NET® is a trademark of USDOL/ETA. People Team AI has modified all or some of this information. USDOL/ETA has not approved, endorsed, or tested these modifications.