Fracture one seductive causal arrow into seven gates, then classify every paper before reaching a verdict.
The driving question
Does AI task exposure predict realized labor-market and macroeconomic outcomes?
AI is already changing tasks and some workplaces, but capability becomes economic consequence only through exposure, adoption, use, organizational response, and equilibrium adjustment.
The seductive claim
You are the juror
“AI can do the task. Therefore the job disappears.”
AI capability→Job loss
Every skipped gate is an empirical question.
Paper-by-gate evidence matrix
No paper spans the whole causal chain.
Direct / observedModeled / indirectNot established
Paper
Capability
Exposure
Adoption
Use
Task productivity
Organization
Labor / macro
2015David H. Autor
absent
absent
absent
absent
absent
modeled
modeled
2018Daron Acemoglu and Pascual Restrepo
modeled
modeled
modeled
modeled
modeled
modeled
modeled
2020Daron Acemoglu and Pascual Restrepo
absent
direct
direct
direct
absent
absent
direct
2021Erik Brynjolfsson, Daniel Rock, and Chad Syverson
absent
absent
modeled
modeled
modeled
modeled
modeled
2023Shakked Noy and Whitney Zhang
direct
direct
direct
direct
direct
absent
absent
2024Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock
direct
direct
absent
absent
absent
absent
absent
2024Tania Babina, Anastassia Fedyk, Alex He, and James Hodson
absent
absent
direct
modeled
absent
direct
direct
2025Daron Acemoglu
modeled
modeled
modeled
modeled
modeled
modeled
modeled
2025Erik Brynjolfsson, Danielle Li, and Lindsey Raymond
direct
direct
direct
direct
direct
direct
absent
2025Anders Humlum and Emilie Vestergaard
absent
direct
direct
direct
modeled
direct
direct
One queue, three futures
The same 15% gain can end in different places.
Only the first move is evidence-anchored. Everything after organizational choice is an illustrative scenario.
+15%resolved issues per hour
ObservedOrganizational choiceIllustrative
Output index111
Staffing index99
Worker-time index98
New-task share6%
With responsive demand, most saved time becomes additional service rather than fewer workers.
Failure → repair
Exposure, adoption, use, productivity, organizational response, and realized outcomes are empirically separate layers.
Observed estimates remain distinct from exposure measures, theory, calibration, and explicitly illustrative scenarios.
The canon · Ten landmark works
Read the contribution. Then read the boundary.
Chronological, not ranked. Influence records intellectual reach—not endorsement or empirical validation.
012015
David H. Autor
Why Are There Still So Many Jobs? The History and Future of Workplace Automation
Theory / synthesis
What it made visible
Task-level comparative advantage and human–machine complementarity
Why it mattered
It replaced occupation-extinction forecasts with a task lens: technology substitutes for some activities while complementing workers in others. It also showed why stable aggregate employment can coexist with polarization and unequal gains.
The limit
It does not identify which new tasks will emerge, how quickly they will scale, or who will be able to enter them.
The Race between Man and Machine: Implications of Technology for Growth, Factor Shares, and Employment
Theory
What it made visible
Automation's displacement effect versus the reinstatement effect of new tasks
Why it mattered
It formalized active capital as an expanding task frontier and made the direction of innovation endogenous. Productivity, wages, employment, labor shares, and inequality can therefore move in different directions.
The limit
The model does not reveal which institutions will induce human-complementary tasks rather than low-value displacement.
Geographically uneven exposure to actual industrial-robot adoption
Why it mattered
It showed that automation can impose substantial local displacement even when national aggregates look calm. One additional robot per thousand workers was associated with a 0.2-point lower employment-to-population ratio and 0.42% lower wages in the aggregate estimates.
The limit
Industrial-robot estimates do not transfer mechanically to generative AI, and long-run migration and occupational adjustment remain uncertain.
The Productivity J-Curve: How Intangibles Complement General Purpose Technologies
Theory + accounting / measurement exercise
What it made visible
Complementary intangible capital and the productivity J-curve
Why it mattered
It explained why transformative technologies can initially coincide with weak measured productivity as firms build unmeasured workflows, skills, data, and organizational capital. Buying a model is not the same as completing the transformation.
The limit
The size and timing of AI-specific complementary investment remain poorly measured, so the J-curve can explain delay without proving a future boom.
Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence
Randomized experiment
What it made visible
Generative AI as a task-level productivity shock for professional writing
Why it mattered
Among 453 professionals, ChatGPT reduced completion time by roughly 40% and raised evaluator-rated quality by about 18%. Lower initial performers benefited more, compressing within-task productivity inequality.
The limit
The experiment measures bounded writing tasks, not fact-checking-intensive workflows, employment, wages, or economy-wide productivity.
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock
GPTs are GPTs: Labor Market Impact Potential of LLMs
Exposure / classification
What it made visible
LLM task exposure and complementary software as a general-purpose-technology multiplier
Why it mattered
It established a reproducible framework for mapping LLM capabilities onto occupational tasks and showed how software built around a model can expand the technical frontier. It became the canonical exposure map for generative AI.
The limit
Exposure is a possibility set, not observed adoption, automation, displacement, wage change, or job loss.
Tania Babina, Anastassia Fedyk, Alex He, and James Hodson
Artificial Intelligence, Firm Growth, and Product Innovation
Observational / IV
What it made visible
AI-skilled human capital as firm investment and product-innovation capital
Why it mattered
AI-investing firms subsequently grew faster in sales, employment, and valuation, principally through product innovation. The gains concentrated among larger firms and coincided with greater industry concentration.
The limit
Causal interpretation still depends on the instrument, and the evidence does not settle whether smaller firms will diffuse AI or fall further behind.
Hulten-style aggregation of affected task shares and task-level cost savings
Why it mattered
It imposed accounting discipline on spectacular macro forecasts and estimated no more than a 0.66% TFP increase over ten years from then-available evidence. It also exposed the danger of extrapolating from easy, measurable tasks to hard, context-dependent work.
The limit
The calibration may miss future capability jumps, AI-enabled innovation, demand effects, new tasks, and adoption frictions in either direction.
Erik Brynjolfsson, Danielle Li, and Lindsey Raymond
Generative AI at Work
Quasi-experimental field evidence
What it made visible
AI as a mechanism for capturing and diffusing tacit expert practice
Why it mattered
Across 5,172 support agents, the assistant increased resolved issues per hour by 15% on average, with much larger gains for less-experienced and lower-skilled workers. It showed how AI can compress performance gaps inside a firm.
The limit
The evidence comes from one firm and a bounded support environment; long-run staffing, expert learning, quality, and wage effects remain open.
Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI
Observational + quasi-experimental
What it made visible
Task reorganization can precede measurable changes in earnings and hours
Why it mattered
Across 11 exposed Danish occupations, about 25,000 workers, and 7,000 workplaces, chatbot initiatives, reported time savings, and new AI tasks spread rapidly while administrative data showed precise null effects on earnings and hours. The March 2026 revision rules out average effects larger than 2% two years after launch—a strong guardrail against converting task gains into instant labor-market claims.
The limit
The evidence covers an early horizon and selected Danish occupations; effects may emerge later, outside earnings and hours, or through margins the administrative data do not capture.
These ten works are ordered by first publication. Selection weighs paradigm effect, conceptual durability, downstream reach, cross-generational influence, and non-redundancy. Every entry names a contribution and a boundary.