The Future of Work Debate Has an Evidence Problem
The most influential number in the AI-and-work debate is also one of the most misunderstood. When researchers at OpenAI published GPTs are GPTs in 2023, their headline finding — that 80% of U.S. workers have at least 10% of their tasks exposed to large language models — rippled through global policy circles. The IMF cited it. The European Parliament cited it. It appeared in U.S. Senate proposals and think tank briefs across multiple continents.
Cohere Labs, in a new analysis published yesterday, does something unusual: instead of adding another number to the pile, they trace where that number went and ask hard questions about whether the evidence driving policy is fit for purpose.
What the scores actually measure
The 80% figure is not a forecast. It's a measurement of technical feasibility: can a GPT-4-era model complete a specific task faster with AI assistance? The answer is evaluated against the U.S. Department of Labor's O*NET taxonomy, which decomposes every occupation into discrete, verifiable tasks.
That specificity carries three compounding limitations:
- Model vintage. The scores reflect a model from early 2023. Frontier capabilities have improved substantially since — one index estimates a roughly 26 percentage point gap between GPT-4-era and current AI capabilities.
- Geographic scope. O*NET is an American taxonomy. It does not transfer cleanly to other labor markets, even with translation. Yet the scores are being used to inform policy in the UK, EU, and beyond.
- Task decomposition. Modeling work as a bundle of scorable tasks captures what can be itemized — not the judgment, relationships, and context that constitute the most consequential parts of most jobs.
These are acknowledged limitations in the original paper. The problem is they travel poorly. When a score calculated against a 2023 model using an American task taxonomy ends up driving retraining budgets in Germany, the limitations don't cancel out — they compound.
The research community is responding
The analysis highlights a growing body of work that addresses these gaps directly:
- Dynamic indexes evaluate AI capabilities as they exist today rather than in 2023. One recent study finds that a 10-point increase in dynamically-measured exposure predicts a 5.6–8.5 percentage point decline in employment — the first empirical evidence that exposure scores predict actual outcomes, not just theoretical susceptibility.
- Ensemble approaches combine multiple exposure frameworks. Individual scores turn out to be weakly or even negatively correlated with each other — they're capturing different dimensions entirely.
- Task-framework extensions examine task adjacency within jobs. The sequencing of AI-exposed tasks changes which occupations appear most at risk.
- Worker-centered measures add what everything else leaves out: what workers actually want. One study finds a substantial category of tasks that AI could perform but that workers do not want automated.
What's missing from the frame
The most striking absence is data workers themselves — the people who label, rate, and curate the training data that powers every LLM whose capabilities are being evaluated against O*NET tasks. They are structurally embedded in the system yet remain invisible in the policy debate those scores enable. A recent ILO review of AI exposure research concluded plainly: the most widely used indicators tell us something meaningful, but not everything we need to know about who is at risk and why.
Why it matters
For anyone building or deploying AI systems, this analysis is a reminder that the policy conversation is running on stale inputs. The 80% number is not a mandate — it's a snapshot of what one model could do in early 2023 under a specific set of assumptions. Three years later, the gap between that snapshot and the decisions it's being used to support has widened to the point of distortion.
The report's recommendations are sober and practical: policymakers should treat exposure scores as one signal among several and invest in interventions whose value doesn't depend on any single forecast being correct. Researchers should build measurement tools that update alongside AI capabilities and extend beyond U.S. labor markets.
The future of work debate is asking three distinct questions — whether AI will advance, what that means for economic outcomes, and what to do about it. The evidence base for the first question is strong. The other two remain far thinner than the policy conversation assumes.