Anthropic

Anthropic Measures the Gap Between AI Ability and Real Use

Anthropic's new 'observed exposure' paper weighs real Claude usage against capability: programmers reach 75% task coverage, while 33% actual use trails 94% feasibility in computer and math work.

Anthropic Measures the Gap Between AI Ability and Real Use — article cover
On this page6 SECTIONS
  1. What “Observed Exposure” Means
  2. Capable Isn’t the Same as Used: The 3x Gap
  3. The Exposure Rankings and the 30% at Zero
  4. Early Signals: No Unemployment Rise, Slower Youth Hiring
  5. How to Read This Study
  6. Sources

On March 5, 2026, Anthropic published “Labor market impacts of AI: A new measure and early evidence,” authored by Maxim Massenkoff and Peter McCrory. The paper introduces a new metric called “observed exposure”: instead of asking what AI could do, it asks what AI is actually being used to do. In a year when everyone in the opening outlook for 2026 expected agents to take over whole workflows, this is a rare labor-market measurement built on real usage data rather than capability demos.

The headline finding fits in one sentence: there is roughly a threefold gap between theoretical capability and actual use.

What “Observed Exposure” Means

Most prior research on AI and jobs stopped at “theoretical exposure”: take a task inventory, ask whether a model can do the task, and score accordingly. That approach has long been criticized for overestimating impact, because capable is not the same as used. Anthropic stacks three data sources instead: the O*NET database’s task breakdowns for roughly 800 US occupations, real Claude conversation data from the Anthropic Economic Index, and the task-level exposure scores from Eloundou et al. (2023).

The scoring rules reflect the intent: fully automated uses of a task get full weight, augmentative uses get half, and work-related usage counts more than personal usage. Task-level coverage is then aggregated to occupations, weighted by time spent on each task. In short, this is an index of automation that has already happened — not automation that might.

Capable Isn’t the Same as Used: The 3x Gap

The most persuasive number in the paper comes from computer and math occupations: 94% of their tasks are theoretically covered by LLM capability, but actual usage covers only 33% — a near-threefold gap. Overall, 97% of Claude usage falls into theoretically feasible task categories, while actual coverage varies enormously across occupations.

That gap is the pressure gauge for the labor market over the next few years. Model capability has already run ahead of adoption; when the impact arrives depends on how fast the gap closes, not on what the next model can do. For product teams it is also a quantified opportunity: two-thirds of the computer-related tasks that could theoretically be automated are not being caught by any product today.

The Exposure Rankings and the 30% at Zero

CBS News and Investopedia both picked up the leaderboard: computer programmers top the list at roughly 75% task coverage, followed by customer service representatives at about 70% and data entry keyers at 67%. On the other end, about 30% of US workers have zero exposure — cooks, bartenders, dishwashers, lifeguards, motorcycle mechanics, dressing room attendants. AI’s reach is heavily concentrated in work that happens through a screen.

The profile of the exposed group overturns a few assumptions: highly exposed workers are 16 percentage points more likely to be female, earn 47% more, and hold graduate degrees at a 17.4% rate versus 4.5% for the unexposed. AI is moving through high-education, higher-paid white-collar work first — not the reverse. Two more demographic numbers: exposed workers are nearly twice as likely to be Asian and 11 percentage points more likely to be white. Methodologically, the authors run a difference-in-differences design on the Current Population Survey, comparing workers in the top exposure quartile against the roughly 30% with zero exposure.

Early Signals: No Unemployment Rise, Slower Youth Hiring

The bad news first: so far there isn’t any. Since ChatGPT’s launch in late 2022, unemployment for highly exposed workers has not risen systematically — the research design could have detected a differential increase of roughly 1 percentage point, and found none.

But two directional signals stand out. First, job-finding rates for workers aged 22–25 entering exposed occupations fell by about 14% after ChatGPT — a change the authors describe as “just barely statistically significant,” with no matching decline for workers over 25. This lines up with Brynjolfsson et al. (2025), who used ADP payroll data and found a 6–16% employment decline for ages 22–25 in exposed occupations. Put differently, the current evidence looks less like “incumbents being replaced” and more like “the entrance narrowing” — new workers can’t get in while the people already seated stay put. Second, in BLS ten-year projections, every 10-percentage-point increase in observed exposure corresponds to a 0.6-point drop in projected 2024–2034 job growth — and running the same regression on conventional theoretical-exposure scores shows nothing at all.

How to Read This Study

Keep three limitations in mind. First, the usage data comes only from Claude: heavy white-collar and developer populations are naturally overrepresented, so this measures automation among Claude users, not all AI automation. Second, exposure is not displacement — within that 33% coverage, most usage is still augmentative, and fully automated usage is a smaller share. Third, usage data necessarily lags frontier capability: the paper measures deployments that have already happened while model capability advances every quarter, so true exposure is more likely understated than overstated.

Still, as the first framework that separates capability from usage and connects directly to official labor statistics, this study hands policymakers and product teams a dashboard worth tracking. What to watch next is not the unemployment rate — it is how fast that 33% closes toward 94%, and whether the job-finding curve for 22-to-25-year-olds keeps bending down.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL