AI productivity is no longer only a conversion problem. Interface design decides who can turn general capability into a verified economic action.
August moves the AI economics frontier from implementation to allocation: interfaces distribute attention, judgment, permission, and action, determining who converts general capability into economic consequence.
2026-08-01 to 2026-08-31 12 ranked papers Cutoff 2026-09-01
01
Executive signal
The allocation layer moved into view
August's strongest AI economics did not ask whether the models had become more capable. It asked what happened between capability and consequence.
That interval now has enough evidence to deserve its own economics. A nationally representative task survey finds generative AI across 80% of occupations and 40% of job tasks, yet adoption remains below 50% in most of them. Exposure scores explain only about half of worker-level variation. In Tennessee middle schools, 96% of students tried an AI tutor, but the median student messaged it on only one-third of practice days and in 17% of sessions containing a mistake. A different tutoring experiment made help appear at precisely those mistakes and paired it with a mastery rule. Students slowed down, recovered more effectively, and produced a modest delayed-learning signal. In Dallas County, an AI tax assistant achieved 78% take-up and raised self-filed property-tax appeals by 9.1 percentage points—but less-advantaged households benefited less.
These are not three stories about “AI access.” They are three different interfaces, each allocating attention, judgment, and action differently. The interface has become an economic institution.
The labor-market evidence also became more granular. One revised payroll study finds no economy-wide displacement but reports that employment among 22–25-year-olds in AI-exposed occupations sits 19% below the path implied by less-exposed peers. The gap operates mainly through hiring, not separations, and attenuates under some controls. It is an important warning, not a causal verdict. A separate firm-level paper associates recent AI investment with productivity growth only where AI-skilled workers appear to build organization capital. That mechanism fits the experimental evidence, but the estimate remains observational.
The broader economics of August tells the same story in older settings. Manager incentives can block internal talent mobility. Noncompetes constrain worker movement without buying measurable secrecy. Firm wage norms can make a temporary inflation burst leave four-year real-wage scars. Rate-of-return regulation can preserve coal capacity long after relative costs change. Institutions do not merely slow adjustment; they decide which margins adjust.
The intellectual shift is therefore from implementation to allocation. July showed that complements convert capability into output. August shows that the design of those complements distributes agency. The scarce asset is no longer only organizational capital. It is the ability to translate a model's general competence into a timely, verifiable, consequential action.
01B
The month in economics
Seven mechanisms. One allocation problem.
Each record separates the development, its consequence, and the confidence the design can support.
01
The interface became an economic institution
Three August experiments identify different conversion margins. In Taxpayer Behavior in the Age of AI, adding an AI chatbot to an already personalized tax-appeal website raises filing without a human agent from 41.4% to 50.5%; 78% of treated households start a conversation. In One Click Away, assignment to Khanmigo raises mathematics achievement by 1.3 national percentile ranks per term, but active engagement is thin. In Making AI Tutoring Productive, an embedded tutor and a mastery rule create “productive friction”: slower progress, better post-error recovery, and a small delayed-test gain on practiced material.
Why it matters
Access, take-up, productive use, and outcome are different variables. A chat window that waits for demand, an assistant that intervenes at an error, and an agent that helps execute a high-stakes administrative act are economically different technologies even if the underlying model is similar.
Strong within the three settings; moderate externally. The designs are randomized. The settings—property-tax appeals and middle-school mathematics—are specific, and the education outcomes remain modest or short-horizon.
02
Exposure is giving way to observed adoption—and to within-occupation variation
Bick, Blandin, Deming, and Schumacher link worker-reported generative-AI use to detailed ONET tasks. At least one in five workers uses generative AI in 80% of occupations and 40% of job tasks, but most occupation-task cells remain below 50% adoption. Exposure measures explain only about half the cross-worker variation; chat-log classifications overstate generic tasks relative to worker surveys.
Why it matters
Occupation-level exposure increasingly looks like a prior rather than a treatment. If workers performing similar work adopt systematically differently, the relevant economic object is the joint distribution of task, worker, firm, and workflow—not an occupation score.
Moderate-to-strong for measurement; early for outcomes. The survey is nationally representative and the taxonomy is transparent. Self-reported use can still be noisy, and observed adoption does not identify causal labor effects.
03
AI labor adjustment may begin at the entry gate
The 12 August revision of Canaries in the Coal Mine? extends ADP payroll data through June 2026. It finds no widespread economy-wide displacement. It does find that employment among workers aged 22–25 in AI-exposed occupations is 19% below the path implied by less-exposed peers, with the divergence concentrated in reduced hiring and substitutive uses.
Why it matters
A technology can leave aggregate employment stable while narrowing the first rung of a career ladder. That distinction matters for human-capital formation: entry-level work is not only output; it is how tacit knowledge and future senior workers are produced.
Early-to-moderate. The administrative sample is large and the pattern survives several alternative controls. The authors explicitly call it descriptive: education controls attenuate it, some divergence predates generative AI, and the ADP pattern is larger than national-survey benchmarks.
04
Organization capital is the proposed bridge from AI investment to firm productivity
Babina, He, and Jiang construct a firm-level AI-investment measure from AI-skilled employment and a new organization-capital measure from job descriptions. AI investment is associated with productivity growth in recent years but not over the previous decade; the association is driven by AI-skilled jobs that build firm-specific knowledge.
Why it matters
The paper gives empirical content to a familiar intangible-capital story. The return to an AI hire may arrive through codified processes, data pipelines, evaluation routines, and redesigned handoffs rather than the worker's immediate output.
Emerging. The timing and job-description mechanism are suggestive, but adopting firms select into AI and organization-capital accumulation. The abstract does not establish a causal productivity coefficient, so none is manufactured here.
05
Cheap expertise can widen inequality after access is equalized
The tax-assistance experiment finds smaller filing gains among less-advantaged households. Choukhmane, de Silva, Lin, and Akuzawa find that AI financial advice broadly moves simulated households toward life-cycle prescriptions, but prompt differences and model responses compound into 4–5% retirement-wealth gaps between groups. For gender, two-thirds of the equity-advice difference comes from men and women writing different prompts; one-third comes from different recommendations for otherwise identical prompts carrying randomized gender labels. The World Development Report 2026 similarly argues that local infrastructure, language, skills, and institutions determine whether rapid diffusion becomes development.
Why it matters
Equal model access can lower the price of expertise without equalizing the capacity to elicit, interpret, trust, or act on advice. Distribution can therefore reappear on the demand side after the supply price collapses.
Strong for the tax treatment effect; moderate for advice content; early for lifetime distribution. The retirement results are simulations under compliance, not observed household wealth.
06
Labor-market frictions leave measurable rents and scars
In the accepted QJE article Clause and Effect, removing a noncompete raises mobility between two competing finance employers by 36–52% and total earnings from them by 12–17%, without detectable secret leakage. In Sticky Wage Norms, 43% of workers staying with one firm from 2021 through 2024 suffer a real-wage decline; even after including job changers, 37% of workers lose ground.
Why it matters
Contracts and norms can matter more than textbook price adjustment. Worker mobility is both an escape valve and a scarce resource: it counters noncompetes and sticky raises, yet most workers cannot or do not switch often enough.
Strong. The noncompete study is a preregistered field experiment; the wage study uses payroll records covering roughly 16 million US workers monthly. The latter's wage-setting mechanism is exceptionally well measured, though its sentiment interpretation is less causally isolated.
07
Institutions decide which margin bears adjustment
Haegele finds that three-quarters of managers report talent hoarding and that quasi-random relief from hoarding raises internal applications. Gowrisankaran, Langer, and Reguant estimate that a regulated utility facing a carbon tax cuts short-run coal generation only 48% as much as a cost minimizer; after 30 years, the cost minimizer has retired 71% more coal capacity. Vivalt et al. find that a three-year guaranteed income reduces labor-force participation by 4.2 percentage points and work by one to two hours weekly, with leisure—not better job quality or degree completion—the largest alternative use.
Why it matters
“Adjustment” is not one outcome. Manager incentives shift it into blocked promotion, utility regulation into legacy capacity, and unconditional income into time. Institutional design determines the margin before prices or technology determine the magnitude.
Strong in the measured settings. The three studies use different designs—quasi-random personnel variation, structural estimation, and a large RCT—so their common lesson is conceptual rather than a pooled causal claim.
02
Economics Radar
An evidence field, not a news feed
August supplies evidence status and direction without manufacturing a numeric rank where the manuscript provides none.
Methods, institutions & behavior9 records · mean score 86
Labor & distribution5 records · mean score 92
Macro & growth3 records · mean score 88
Climate & resources2 records · mean score 91
Firms & industrial change2 records · mean score 87
Paper skylineRadar score by rank
0194
0293
0392
0490
0589
0688
0787
0893
0994
1092
1193
1291
04
The papers that matter
Twelve contributions worth carrying forward
Each record separates status, method, result, relevance, and limitation.
0194
Strong within the study setting; external validity remains boundedNBER Working Paper 35632, August 2026; DOI 10.3386/w35632.
Taxpayer Behavior in the Age of AI: A Field Experiment on Property Tax Appeals
Justin E. Holz (University of Michigan), Ricardo Perez-Truglia (UCLA), Andrew Simon (University of Virginia), and Alejandro Zentner (University of Texas at Dallas).
Chatbot access raises filing without a human agent by 9.1 percentage points, from 41.4% to 50.5%. 78% of treated households initiate a conversation. Clickstream and transcript evidence points to assistance with judgment rather than mere information retrieval.
Method, implication, and boundary
Question
Can a low-cost AI assistant help households convert tax information into an appeal, and who benefits?
Method / data
Field experiment with 645 Dallas County households. Everyone received personalized assessment information, instructions, and evidence; half also received a tailored AI chatbot.
Why it matters
It is unusually clean evidence that AI can reduce the cost of expert administrative action after basic information has already been supplied.
Limitations
One county, one tax process, and a self-selected study sample. The treatment effect is smaller for less-advantaged households, but the study describes that heterogeneity as suggestive rather than definitive.
Moderate — serious working-paper or institutional evidenceNBER Working Paper 35677 and CEPR Discussion Paper 21889, August 2026; CEPR issue date 31 August.
What Work Does Generative AI Do?
Alexander Bick (Federal Reserve Bank of St. Louis), Adam Blandin (Vanderbilt University), David J. Deming (Harvard University), and Tyler R. Schumacher (Vanderbilt University).
At least 20% of workers use generative AI in 80% of occupations and 40% of job tasks, but adoption stays below 50% in most. Exposure measures explain only about half the variation across workers. Chat-log measures disproportionately classify conversations into generic tasks.
Method, implication, and boundary
Question
Which workers use generative AI for which tasks, and how does realized use differ from technical exposure or platform-chat classifications?
Method / data
Nationally representative worker survey linked to detailed occupations and ONET tasks; construction of task- and occupation-level adoption indices.
Why it matters
It changes the empirical denominator. Labor research can now study observed task adoption instead of treating capability exposure as actual use.
Limitations
Survey reporting and task matching introduce measurement error; the indices describe adoption but do not identify employment, wage, or productivity effects.
Strong within the study setting; external validity remains boundedNBER Working Paper 35620, August 2026; EdWorkingPaper 26-1551.
One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment
Philip Oreopoulos and Nina Low (University of Toronto).
Assignment raises mathematics achievement by 1.3 national percentile ranks per term, approximately 0.06–0.08 standard deviations over a school year. A full year of active participation implies about 0.14 standard deviations. Yet 96% try the tutor at least once while the median student messages it on only one-third of practice days and in only 17% of mistake-containing exercise sessions.
Method, implication, and boundary
Question
Does access to a guard-railed AI tutor improve mathematics achievement at scale?
Method / data
Two-year cluster-randomized trial in 18 Tennessee middle schools. Assigned grades used Khan Academy with Khanmigo during daily remedial mathematics blocks.
Why it matters
It isolates engagement as a binding input. A technically available tutor produces gains similar to Khan Academy practice without AI when students rarely engage it at the moment of need.
Limitations
The estimate is assignment-to-access, not compulsory use; grades are randomized within schools; the active-participation estimate is not the same as the randomized intent-to-treat effect.
Strong within the study setting; external validity remains boundedEdWorkingPaper 26-1552, 18 August 2026; NBER Working Paper 35621; DOI 10.26300/01qv-6c22.
Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment
Philip Oreopoulos and Michael Liut (University of Toronto), Alp Sungu (University of Pennsylvania), and Nina Low (University of Toronto).
Mastery adds 4.62 attempted questions, 1.23 correct answers, and 28.7 percentage points to the probability of reaching three correct answers in a row on the initial exercise, but does not alone improve delayed learning. AI reduces completed questions by about one and adds 1.64 minutes per question while improving post-error recovery. Inside the mastery condition, AI raises the practiced Exercise 1 delayed-test outcome by about 3 percentage points; the estimate is only marginally significant.
Method, implication, and boundary
Question
Can a mastery workflow make students use AI support more productively after mistakes?
Method / data
Individual randomization of more than 6,000 middle-school students to AI versus standard computer-assisted learning, mastery versus non-mastery progression, and one of two mathematics topics.
Why it matters
It treats interface timing and incentives as components of the intervention. The result is not “AI works”; it is that structured support can make some slowness productive.
Limitations
A one-week delayed test, one short intervention, and a positive result concentrated in one practiced exercise. The paper explicitly rejects a transformational interpretation.
Moderate — serious working-paper or institutional evidenceStanford Digital Economy Lab working paper, originally 2025; materially revised 12 August 2026 with payroll data through June 2026.
Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence
Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen (Stanford Digital Economy Lab; Brynjolfsson and Chen also Stanford HAI).
No evidence of broad economy-wide displacement. Employment among workers aged 22–25 in exposed occupations is 19% below the trajectory implied by less-exposed peers; the gap is concentrated in hiring, substitutive uses, and employment rather than base pay.
Method, implication, and boundary
Question
Are early labor-market changes concentrated among particular workers and kinds of AI use?
Method / data
High-frequency ADP administrative payroll data covering millions of US workers; exposure and use classifications; robustness checks excluding technology firms and computer occupations and controlling for remote work and interest-rate exposure.
Why it matters
It makes entry-level hiring the most plausible near-term pressure point in the US aggregate data.
Limitations
Descriptive, not causal. Education controls attenuate the result, some divergence predates generative AI, and the ADP estimate exceeds national-survey benchmarks.
Moderate — serious working-paper or institutional evidenceNBER Working Paper 35684 and CEPR Discussion Paper 21894, 31 August 2026.
Canaries in the Gold Mine: Early Productivity Gains from Artificial Intelligence Creating Organization Capital
Tania Babina, Alex X. He, and Renhao Jiang (University of Maryland, Robert H. Smith School of Business).
AI investment is associated with productivity growth in recent years but not in the prior decade. The recent association is driven by AI-skilled jobs whose descriptions indicate accumulation of durable firm-specific knowledge.
Method, implication, and boundary
Question
Are recent firm productivity gains associated with AI investment, and what intangible asset mediates them?
Method / data
A firm-level AI-investment measure based on AI-skilled employment from machine learning through generative and agentic AI; job descriptions used to construct a measure of organization-capital-building work.
Why it matters
It offers a measurable bridge between AI hiring and productivity and turns “organizational complementarity” into a testable labor-composition mechanism.
Limitations
Firm selection, reverse causality, and generated measures remain central. The public abstract provides no headline coefficient suitable for reproduction, and the result should not be written as a causal productivity effect.
Moderate — serious working-paper or institutional evidencearXiv:2608.01607, submitted 3 August 2026; NBER Working Paper 35574.
AI Financial Advice: Supply, Demand, and Life Cycle Implications
Taha Choukhmane (MIT Sloan and NBER), Tim de Silva (Stanford Graduate School of Business and Stanford HAI), Weidong Lin and Matthew Akuzawa (MIT Sloan).
Following model advice produces diversified equity participation for more than 99% of simulated individuals, declining equity shares after age 45, and larger buffers. It also relies too heavily on simple heuristics and smooths consumption poorly after job loss. Group-specific prompts and responses compound into 4–5% retirement-wealth differences.
Method, implication, and boundary
Question
What advice do people elicit from frontier models, and how would following it compound over a lifetime?
Method / data
Prompts from a demographically representative sample of 1,000 US adults; repeated GPT-5.2 and Gemini 3 Flash advice; calibrated life-cycle simulations with income, unemployment, asset-return, tax, and mortality risk.
Why it matters
It separates the supply of advice from the demand embodied in a prompt, revealing a durable distributional channel that model improvement alone may not remove.
Limitations
Advice is not observed behavior; welfare depends on the calibrated environment and assumed compliance; model versions will change; Prolific quotas are not a probability sample.
Moderate — serious working-paper or institutional evidenceNBER Working Paper 35624 and BFI Working Paper, 12 August 2026.
Sticky Wage Norms and the Real Wage Cost of Unexpected Inflation
Erik Hurst and Christina Patterson (University of Chicago Booth and NBER), Nela Thomas Richardson and Ye Liv Wang (ADP Research).
43% of workers continuously at one firm over 2021–2024 have lower real wages after four years, with an average loss of roughly 9% among those losing ground. Including job changers leaves 37% below their starting real wage. One-for-one indexation of modal raises would close roughly 40% of the shortfall from the pre-pandemic trend.
Method, implication, and boundary
Question
Why did nominal wages fail to absorb the 2021–2024 inflation shock for so many workers?
Method / data
ADP payroll records covering roughly 16 million US workers per month, 2016–2025; firm-level modal raise rules, job-stayer and job-changer comparisons, counterfactual indexation, and cross-country evidence.
Why it matters
It turns nominal rigidity from an aggregate parameter into an organizational norm—often a 3% default raise applied across a firm.
Limitations
The payroll evidence is descriptive; the counterfactual holds other behavior fixed; Belgian indexation is informative but not a randomized test of sentiment.
Strong — peer-reviewed or accepted researchAccepted manuscript, Quarterly Journal of Economics, published online 7 August 2026; DOI 10.1093/qje/qjag040.
Clause and Effect: Theory and Field Experimental Evidence on Noncompete Clauses
Bo Cowgill (University of Toronto), Brandon Freiberg (INSEAD), and Evan Starr (University of Maryland).
Removing the clause raises mobility between the firms by 36–52% and total earnings from them by 12–17%. The experiment rejects even small secret-leakage effects and finds no wage premium for accepting the restraint.
Method, implication, and boundary
Question
Do noncompetes protect secrets or exploit worker inattention and uncertainty about enforcement?
Method / data
Preregistered field experiment across approximately 14,000 job offers to freelance recruiters at two finance firms; wages and noncompete presence, salience, and duration randomized.
Why it matters
It supplies unusually direct causal evidence that a common labor contract can create monopsony-like frictions without delivering its stated informational benefit.
Limitations
Short-term freelance contracts at two firms; secrecy is measured in this setting; longer-horizon training and client investment may differ elsewhere.
Strong — peer-reviewed or accepted researchAmerican Economic Review 116(8), August 2026, 3110–3151; DOI 10.1257/aer.20220264.
Talent Hoarding in Organizations
Ingrid Haegele (Ludwig-Maximilians-Universität München).
Three-quarters of managers acknowledge hoarding. The behavior is visible in ratings and increases with performance pay, team size, and talent visibility. Relief from hoarding raises internal applications and changes who advances.
Method, implication, and boundary
Question
Do managers suppress the internal mobility of talented workers when team-level incentives make departures personally costly?
Method / data
Personnel records and manager surveys from a European manufacturer with more than 200,000 employees; manager rotations provide quasi-random variation in workers' exposure to hoarding incentives.
Why it matters
Firms can possess the relevant talent and still misallocate it. That is a direct warning for AI strategies built around a small pool of technical employees and business-unit managers evaluated on local output.
Limitations
One large organization; rotations may not eliminate every correlated managerial change; self-reports can be strategic.
Strong — peer-reviewed or accepted researchAccepted manuscript, Quarterly Journal of Economics, published online 14 August 2026; DOI 10.1093/qje/qjag042.
The Employment Effects of a Guaranteed Income: Experimental Evidence from Two U.S. States
Eva Vivalt (University of Toronto), Elizabeth Rhodes and Patrick Krause (OpenResearch), Alexander Bartik (University of Illinois Urbana-Champaign), David Broockman (University of California, Berkeley), and Sarah Miller (University of Michigan Ross School of Business).
Non-transfer individual income falls about $1,900 annually, labor-force participation falls 4.2 percentage points, and work falls one to two hours weekly; partners reduce work comparably. Leisure rises most. Job quality and degree attainment do not improve detectably; early well-being gains fade.
Method, implication, and boundary
Question
How does a large unconditional income transfer affect employment, job quality, human capital, time use, and well-being?
Method / data
1,000 low-income adults randomized to $1,000 monthly for three years; 2,000 controls received $50 monthly; surveys, administrative records, and phone-app data.
Why it matters
It is a rare long-duration, high-transfer RCT that rules out several popular offsetting channels rather than measuring only labor supply.
Limitations
Two US states and a selected low-income sample; the control group also received money; effects may differ under permanent national financing and general equilibrium.
Strong — peer-reviewed or accepted researchAmerican Economic Review 116(8), August 2026, 2928–2961; DOI 10.1257/aer.20240094.
Energy Transitions in Regulated Markets
Gautam Gowrisankaran (Columbia University), Ashley Langer (University of Arizona), and Mar Reguant (Northwestern University).
A regulated utility facing a carbon tax reduces short-run coal generation only 48% as much as a cost minimizer. Thirty years after a sudden transition, the cost minimizer has retired 71% more coal capacity.
Method, implication, and boundary
Question
How does rate-of-return regulation change utilities' use and retirement of legacy coal capacity during an energy transition?
Method / data
Structural model estimated on US utility generation and capacity, with regulated firms trading off operating cost, reliability, affordability, and returns on installed capital.
Why it matters
A correct relative price can fail to produce the expected transition when the institution rewards ownership of the legacy asset.
Limitations
Structural conclusions depend on model specification and counterfactual policy assumptions; faster retirement can threaten reliability and affordability rather than dominate on every welfare margin.
Availability, invocation, productive use, judgment, permission, action, and durable output are different economic variables.
The conversion chain
A capable model can fail at nine different margins
Select a layer. The instrument preserves the observed object and the failure that prevents it from becoming economic consequence.
01 / 09Capability
Researchers often observe
Benchmark score or task demonstration
Conversion failure
A technically feasible output is treated as economically reliable, lawful, or worth using.
Read the complete nine-layer measurement ledger
CapabilityBenchmark score or task demonstrationA technically feasible output is treated as economically reliable, lawful, or worth using.
ExposureTextual overlap between task descriptions and capabilitiesPotential substitution is mistaken for adoption or a treatment dose.
AdoptionAny reported use during a reference periodA single trial and daily workflow integration receive the same label.
EngagementMessages, sessions, clicksActivity can reflect confusion rather than productive use; passive nonuse may reflect poor timing.
Task outputTime, quantity, or benchmark qualityFaster completion can reduce verification or learning; task gains need not raise valuable output.
Worker outputIndividual production or performanceTeam spillovers, reassignment, monitoring, and changes in task mix are omitted.
Firm productivityAI hiring, patents, value added, or TFPAdopters are selected; intangible investment can initially lower measured productivity; quality is difficult to price.
Aggregate productivitySectoral or national output per hourDiffusion lags, entry and exit, unmeasured free services, and distribution across firms can offset micro gains.
DistributionAccess by demographic or income groupDifferences in prompts, trust, compliance, execution costs, and market power can recreate inequality after access.
Where the literature stood
By July 2026, three propositions had become difficult to dispute. Frontier model capability was advancing faster than measured economic productivity; the relevant production unit was the task rather than the occupation; and complementary investments in data, workflow, training, software, and management determined how much technical capability became useful output. The unresolved question was how those complements actually worked.
Most empirical work still occupied one of two ends of the conversion chain. Benchmark studies measured what a model could do. Exposure studies mapped those capabilities onto tasks and occupations. Neither observed the sequence in between: whether a person invoked the tool, whether the interface appeared at the right moment, whether the answer changed a decision, whether the decision became an action, and whether the action raised durable output. “AI adoption” compressed all of those margins into one indicator.
The historical comparison with electrification and enterprise software remained useful but incomplete. Both required complementary capital and organizational redesign before appearing in aggregate productivity. Generative AI adds a distinct feature: the complement is partly a decision architecture. The same general-purpose model can wait passively in a chat window, interrupt at an error, draft a form, make a recommendation, or execute a transaction. Interface design therefore changes the economic treatment itself.
What changed this month?
August supplied evidence at several previously missing links. The studies do not form one causal chain—their settings and designs differ—but together they make the conversion problem observable.
Representative prompt collection plus model and life-cycle simulations
Advice supply, demand, and distribution
Different prompts and conditional advice can compound into 4–5% simulated retirement-wealth gaps.
The common update is not that “AI works.” It is narrower and more important: the placement, timing, and permissions of the interface determine which economic margin moves. A passive tutor changes optional help-seeking. An embedded tutor changes error recovery. A tax assistant changes the probability that personalized information becomes a filed appeal. A firm-specific AI role may create reusable routines rather than only individual task output.
Strongest evidence
01
Causal effects are strongest at the local conversion margin
The tax and tutoring experiments deserve the greatest confidence for their measured outcomes. Random assignment identifies what the interfaces caused in those settings. The tax study is especially revealing because both groups already received personalized information and evidence. The additional 9.1-point effect is therefore not well described as information provision; it is assistance with interpretation, confidence, and execution. The limitation is external validity, not internal validity.
The education studies are similarly credible but should not be collapsed into a single “AI tutoring effect.” Khanmigo raised achievement by 1.3 percentile ranks per term, while engagement remained low. The NUMI experiment instead manipulated both AI assistance and a mastery constraint. Its strongest immediate effects concern practice behavior—4.62 more attempts, 1.23 more correct answers, and a 28.7-point increase in completing three consecutive correct responses on the initial exercise. The delayed-test improvement for practiced material is about 3 percentage points and only marginally significant. That is a learning signal, not a settled estimate of durable human-capital accumulation.
02
The best adoption evidence replaces an exposure proxy with reported behavior
The nationally representative task survey is the strongest August measurement contribution. It shows why occupation exposure is insufficient: within a nominally exposed occupation, adoption differs across tasks and workers, and conventional exposure measures explain only about half of that variation. The study is descriptive and relies on self-reports, but its data are much closer to actual use than benchmark-to-task mappings or public chat logs.
03
The labor-market result is a warning with explicit causal limits
The 19% young-worker employment gap is economically large and comes from millions of payroll observations. It is not, however, an experimental or quasi-experimental treatment effect. Education controls attenuate it; some pre-trend exists; and national surveys show a smaller pattern. The correct reading is that the first visible adjustment may be reduced hiring into exposed work—not that AI has been shown to cause a 19% employment decline.
04
Firm productivity and lifetime wealth remain mechanism evidence
The organization-capital paper and the financial-advice study supply plausible mechanisms at scales that experiments do not yet reach. Their claims should stay calibrated. AI investment is selected, not randomly assigned. Retirement wealth is simulated under advice-compliance assumptions, not observed. Both studies are valuable because they identify where to look next, not because they close the causal case.
Emerging hypotheses
Translation capital
The August evidence suggests a complement more specific than generic digital skill: translation capital—the ability, permission, context, and workflow required to turn a model's general competence into a reliable action. It can reside in a person who knows what to ask, an interface that recognizes a mistake, a firm routine that validates an output, or an institution that permits the output to trigger a decision. This is an inference from the combined studies, not a variable any one paper estimates.
The career-ladder hypothesis
If substitution begins with junior hiring, the short-run wage bill may fall while the long-run stock of experienced workers deteriorates. Entry-level tasks often produce both current output and future expertise. Firms that automate the former without rebuilding the latter may discover a delayed human-capital constraint. The payroll evidence makes the hypothesis urgent; it does not yet establish the mechanism.
Organization capital is a distinct AI input
AI-skilled employment may matter less through the direct output of specialists than through the routines they leave behind: evaluation systems, data definitions, escalation rules, reusable prompts, and redesigned handoffs. That would explain why adding the same model to two firms can produce different value. Credible identification of this channel is now the research priority.
Agency inequality can survive universal access
When advice becomes cheap, inequality may migrate from access to invocation, interpretation, trust, and execution. Smaller tax-filing gains among less-advantaged households and simulated wealth gaps from different financial prompts are consistent with this account. Whether interface defaults can close those gaps without paternalism is open.
“Small AI” may outperform scale when context is local
The World Development Report 2026 emphasizes adoption, adaptation, and smaller systems suited to local language and infrastructure. The economic conjecture is that a less capable model embedded in a trusted, low-cost, domain-specific process can create more welfare than a frontier model separated from the relevant institution.
Contradictions
Broad reach, shallow intensity. AI appears in 80% of occupations and 40% of tasks under a 20%-of-workers threshold, yet most occupation-task cells remain below 50% adoption. Breadth is not depth.
Aggregate calm, cohort stress. The payroll evidence finds no widespread employment collapse but a pronounced divergence for 22–25-year-olds in exposed occupations. Aggregation can hide a change in who enters.
Universal availability, unequal conversion. A tax assistant is offered at random, yet gains are smaller for less-advantaged households. Equal treatment availability does not imply equal treatment response.
Speed, but not always learning. AI assistance can reduce the number of questions completed and increase time per question while improving recovery from errors. In a learning technology, productive friction may dominate throughput.
Scale, but local bottlenecks. Frontier capability is centralized, while electricity, connectivity, language, institutional trust, and implementation knowledge remain local. The global diffusion curve can steepen while welfare gaps widen.
These are not logical inconsistencies. They are evidence that the outcome depends on the level of observation and the conversion mechanism.
Measurement problems
Layer
What researchers often observe
What can go wrong
Capability
Benchmark score or task demonstration
A technically feasible output is treated as economically reliable, lawful, or worth using.
Exposure
Textual overlap between task descriptions and capabilities
Potential substitution is mistaken for adoption or a treatment dose.
Adoption
Any reported use during a reference period
A single trial and daily workflow integration receive the same label.
Engagement
Messages, sessions, clicks
Activity can reflect confusion rather than productive use; passive nonuse may reflect poor timing.
Task output
Time, quantity, or benchmark quality
Faster completion can reduce verification or learning; task gains need not raise valuable output.
Worker output
Individual production or performance
Team spillovers, reassignment, monitoring, and changes in task mix are omitted.
Firm productivity
AI hiring, patents, value added, or TFP
Adopters are selected; intangible investment can initially lower measured productivity; quality is difficult to price.
Aggregate productivity
Sectoral or national output per hour
Diffusion lags, entry and exit, unmeasured free services, and distribution across firms can offset micro gains.
Distribution
Access by demographic or income group
Differences in prompts, trust, compliance, execution costs, and market power can recreate inequality after access.
Three improvements follow. Measure intensity and workflow position, not only any use. Preserve the denominators at every conversion step. And connect person-level AI telemetry to administrative outcomes without confusing a predictive exposure index with an instrument for adoption.
06
Numbers worth remembering
Ten quantities with their caveats attached
19%
employment among 22–25-year-olds in AI-exposed occupations was below the path implied by less-exposed peers by June 2026 in the preferred ADP specification; descriptive, not causal.
9.1 percentage points
chatbot access raised self-filed property-tax appeals from 41.4% to 50.5% after both groups received personalized information and evidence.
78%
the share of treated Dallas households that initiated a conversation with the tax assistant.
80% / 40%
the shares of occupations / job tasks in which at least one in five workers reported generative-AI use.
17%
the share of Khanmigo sessions containing a student error in which the student sent the tutor a message, despite 96% trying the tool at least once.
About 3 percentage points
the delayed-test gain on material practiced with AI in the mastery-based tutoring experiment; marginally statistically significant.
4.5% versus 14.2%
the World Bank's estimated shares of jobs with automation potential in developing versus high-income economies; these are task potentials, not displacement forecasts.
4–5%
simulated retirement-wealth gaps generated by group differences in financial prompts and model recommendations under the study's advice-compliance assumptions.
43%
the share of same-firm workers with a real-wage decline over 2021–2024 in the ADP payroll study, up from 21% at the beginning of the period.
36–52%
the increase in mobility between two competing employers after noncompete removal in the QJE field experiment; earnings from the two employers rose 12–17%.
06B
Charts worth remembering
Three figures rendered. Two claims held at the data gate.
Every panel preserves native units and denominators. Missing series and uncertainty intervals are treated as evidence boundaries, not invitations to interpolate.
Static evidence summary: Dallas self-filing rises from 41.4% to 50.5% with AI assistance; 78% start a conversation. Khanmigo is tried by 96%, but appears in 17% of error sessions. Developing- and high-income economies show 4.5% versus 14.2% automation potential and 16.2% versus 18.7% augmentation potential. Real-wage declines rise from 21% to 43% among same-firm workers and from 24% to 37% among all workers between 2021 and 2024.
Figure 01 · Mechanism
Access Is Not the Treatment
At which point does an AI interface convert availability into engagement, action, or learning?
Dallas tax appeal41.4%Control self-filed
Dallas tax appeal
0100%
Control self-filed
AI self-filed
Started AI conversation
Khanmigo
0100%
Tried tutor once
Median practice days used
Error sessions with message
NUMI · exact estimates kept on native units; no common magnitude scale
Attempts
+4.62count difference
Correct answers
+1.23count difference
Mastery completion
+28.7 ppprobability difference
Delayed practiced item
≈ +3 ppmarginally significant
Availability, invocation, persistence, error recovery, execution, and durable outcome are distinct margins. “Users with access” is not a treatment dose.
The panels must remain separate. They differ in population, intervention, endpoint, unit, and follow-up. No funnel or pooled conversion rate is defensible.
Why this form—and which alternatives were rejected?
selectedBest analytical fit
Faceted Cleveland dot plot with one self-contained panel per experiment and arrows only where a genuine within-outcome control-to-treatment comparison exists. A formal chart-form consultation scored this fit 66.1/100.
baselineStandard baseline
Horizontal bars, faceted by study and grouped by outcome, also scored 66.1/100. It is familiar but visually encourages comparison across incompatible units.
rejectedNon-standard challenger
Density small multiples, scored 59.9/100. Rejected by the challenger gate: the published headline data are means and treatment effects, not distributions; simulated densities would add false resolution rather than reveal structure.
Figure 02 · Causal boundary
The Employment Shock May Start at the Hiring Gate
Does the aggregate stability of employment conceal a post-generative-AI divergence among young workers in exposed occupations?
Held at the data gate
The headline survives. The chart waits.
The paper's monthly or quarterly employment-index series through June 2026 by age group and AI exposure; event-time estimates and confidence intervals; hiring and separation decompositions; preferred and education-adjusted specifications.
Required before rendering: x-axis = calendar time or event time around the public diffusion of generative AI; y-axis = employment relative to the paper's pre-period normalization or coefficient relative to less-exposed work.
Stable aggregate employment can coexist with a narrowing entry channel. Hiring flows reveal adjustment earlier than total headcount.
Label the chart descriptive, not causal. The estimate changes with education controls, and ADP is not the entire US workforce.
Why this form—and which alternatives were rejected?
selectedBest analytical fit
Indexed line chart plus a compact specification inset, with exposed and less-exposed young-worker paths and shaded 95% intervals. The inset prevents the 19% headline from floating free of sensitivity checks.
baselineStandard baseline
Horizontal coefficient bars for the current gap by age, exposure type, and control set. Cleaner for magnitude, weaker for showing when divergence begins.
rejectedNon-standard challenger
Horizon graph of age-by-exposure deviations. Rejected by the challenger gate: compression would hide the pre-trend and uncertainty that determine the result's credibility.
Figure 03 · Comparison
Lower Automation Exposure Is Not the Same as Greater AI Opportunity
How do automation and augmentation potential differ between developing and high-income economies?
Automation potential4.5%Developing economies
020%
Automation potential
Augmentation potential
Developing economies may face less immediate task substitution while having almost as much scope for augmentation. The binding constraint is conversion capacity, not technical exposure alone.
These are modeled task potentials, not forecasts of jobs lost or productivity gained.
Why this form—and which alternatives were rejected?
selectedBest analytical fit
Two-row connected dot plot. Each row connects developing and high-income economies, revealing a large automation gap but a small augmentation gap.
baselineStandard baseline
Grouped horizontal bars for the four values. Accurate but less efficient at emphasizing the between-group distance within each mechanism.
rejectedNon-standard challenger
Quadrant scatterplot with automation on one axis and augmentation on the other. Rejected by the challenger gate: two aggregate observations cannot support a meaningful spatial pattern; the form would imply more structure than the data contain.
Figure 04 · Uncertainty
Productive Friction in AI Tutoring
Can an AI tutor improve error recovery even when students complete fewer questions and spend longer on each one?
Held at the data gate
The headline survives. The chart waits.
NUMI arm means, treatment coefficients, standard errors, and sample counts for questions attempted, time per question, correct responses after an error, completion of three consecutive correct responses, and delayed-test accuracy, split by mastery condition where preregistered.
Required before rendering: x-axis = treatment effect with 95% confidence interval; y-axis = outcome, grouped as throughput, recovery, mastery, and delayed learning. Use separate facets for counts, minutes, and percentage points.
For learning, faster is not automatically better. A useful assistant can increase time on the bottleneck and still improve the economically relevant outcome.
The delayed practiced-item gain is about 3 percentage points and marginally significant. Do not visually equate immediate practice behavior with durable learning.
Why this form—and which alternatives were rejected?
selectedBest analytical fit
Faceted coefficient plot. It can display the apparent trade-off—roughly one fewer question and +1.64 minutes per question—beside improved error recovery and the small delayed-learning estimate.
baselineStandard baseline
Grouped bars of arm means. Useful for levels, but crowded in a factorial design and less transparent about estimation uncertainty.
rejectedNon-standard challenger
Sankey diagram from mistakes to recovery to mastery. Rejected by the challenger gate: the published outcomes are not conserved person-level flows, so a Sankey would invent transitions and denominators.
Figure 05 · Persistence
Temporary Inflation, Persistent Real-Wage Scars
How did nominal raise norms transmit the 2021–2024 inflation burst into real-wage losses?
Same-firm workers21%2021
050%
Same-firm workers
All workers
Nominal rigidity can be behavioral and institutional even when nominal wages rise. A repeated 3% raise norm turns an inflation shock into a multiyear real loss.
The study measures payroll outcomes with unusual precision, but the counterfactual share attributable specifically to “norms” depends on its decomposition and employer classification.
Why this form—and which alternatives were rejected?
selectedBest analytical fit
Dumbbell chart connecting 2021 and 2024 for each sample, annotated with the 22- and 13-point increases.
baselineStandard baseline
Grouped bars for the two years and two samples. Familiar but less direct about within-sample change.
conditionalNon-standard challenger
Beeswarm of worker-level real-wage changes. Conditionally accepted only with secure microdata access: it would reveal the distribution hidden by the two shares, but cannot be reconstructed from published aggregates. For the Radar, it is rejected.
07
Where economists disagree
A disagreement is useful when evidence can resolve it
01
Is AI already displacing labor?
Position A
The absence of broad aggregate employment decline and the continued expansion of exposed occupations imply that displacement is not yet economically important.
Position B
Aggregate stability conceals a leading-edge shock to young workers, with employment in exposed occupations 19% below a comparison path and the adjustment concentrated in hiring.
Early warning, not causal verdict. The cohort result is too large to ignore and too design-sensitive to translate into an economy-wide displacement estimate.
Resolution: Employer-level adoption timing linked to vacancy, applicant, hiring, task, and payroll data; credible comparison groups; and follow-up of affected cohorts across occupations.
02
Is access the binding constraint, or is workflow design?
Position A
Once a capable assistant is cheap and broadly available, learning and productivity will follow through voluntary experimentation.
Position B
Optional access produces thin and unequal engagement; timely prompts, mastery rules, context, and execution support determine conversion.
Workflow design is the stronger August signal. Access is necessary in all three experiments, but it is not a sufficient statistic for treatment intensity.
Resolution: Multi-arm trials that independently vary model quality, interface timing, default use, human support, and authority to act, with durable output measured after assistance ends.
03
Does organization capital cause AI productivity, or merely accompany strong firms?
Position A
AI-skilled workers build reusable firm knowledge, which is the mechanism converting AI investment into recent productivity growth.
Position B
Productive, well-managed firms both hire AI talent and write sophisticated job descriptions; the measured organization capital may be a marker of selection.
Plausible mechanism, unproven causal channel. It is the best firm-level hypothesis of the month, not yet a return-on-investment parameter.
Resolution: Staggered deployment with precommitted workflows, exogenous supply shocks to AI talent, or randomized implementation support, combined with value-added, quality, and process telemetry.
04
Will cheap AI advice equalize expertise or reproduce inequality?
Position A
Near-zero marginal-cost advice relaxes information and expertise constraints, particularly for households unable to buy professional help.
Position B
Differences in prompting, trust, interpretation, and ability to execute recreate disparities; model responses can also vary with demographic cues.
Average access improves; relative incidence is unresolved. The causal average treatment effect is strong, while the longer-run inequality magnitude is suggestive or simulated.
Resolution: Large trials powered for heterogeneous effects, randomized defaults and explanations, observed compliance, and administrative measures of realized wealth or benefit receipt.
05
Should AI deployment be governed by liability or by staged permission?
Position A
Existing liability and voluntary risk management can make deployers internalize harm without delaying useful innovation.
Position B
When defensive effort and information arrive over time, liability alone cannot implement the efficient release date; staged access can improve both learning and safety.
Theoretical case for complementarity; little comparative empirical evidence. August clarifies the mechanism but does not identify an optimal regime.
Resolution: Cross-sector evidence on incident rates, learning curves, deployment stages, insurance prices, and user benefits under different liability and permission structures.
08
Emerging research frontier
Seven questions the next editions must track
01
The economics of translation capital
Can the capacity to turn general-purpose AI into a verified action be measured as a productive asset distinct from human capital, software, and organization capital? Research needs workflow-level data on model availability, invocation, context supplied, validation, override, and completed economic actions. Randomized implementation support would identify the return to the complement rather than to model access.
02
AI and the production of future experts
Do firms that reduce junior hiring also reduce later supplies of managers, engineers, analysts, and professionals? Answering this requires linked vacancy, hiring, task, training, promotion, and earnings histories over many years, together with employer adoption timing. The key outcome is not only the first job lost but the expertise that is never produced.
03
From seats to treatment intensity
Which usage measure predicts output: active days, task share, suggestions accepted, verified actions, or workflow coverage? Firms and statistical agencies need common intensity measures that preserve task and worker denominators. Designs should compare passive availability with defaults, embedded triggers, and delegated execution.
04
Distribution after access
Which groups benefit once model availability is equalized, and at which conversion step does inequality emerge? Large field experiments should randomize explanations, defaults, local-language support, and human escalation; collect prompts and trust measures; and link them to administrative outcomes rather than intentions.
05
The depreciation rate of organization capital
Are AI-built prompts, evaluation suites, and process knowledge durable assets or rapidly obsolescent code? Accounting and productivity research needs versioned repositories, staff turnover, model changes, and workflow outcomes. The answer determines whether current spending is investment, maintenance, or rent paid to a platform.
06
Small AI and local institutional fit
When does a smaller, cheaper, locally adapted system outperform a frontier model in welfare terms? Multi-country trials should hold the task constant while varying language coverage, latency, connectivity, privacy, model size, and integration with local institutions. Cost per verified outcome—not benchmark score—should be the common metric.
07
Governance by permission layer
What are the welfare effects of allowing an AI system to inform, draft, recommend, or execute? Sectoral deployments create natural variation in permission and liability. Research should measure benefit, error severity, detection, defensive investment, and learning at each stage rather than classify a system simply as released or not released.
09 · The Schymura Take
Who gets to turn an answer into an action?
Inference — not a finding of any single cited paper.
The first serious distributional conflict around AI may not be over wages. It may be over who gets to turn an answer into an action.
That sounds abstract until the August evidence is placed side by side. Dallas homeowners receive the same personalized tax evidence, but only one group receives an assistant that can help carry the argument across the administrative threshold. Students receive access to a capable tutor, yet most do not summon it when they make a mistake. Another interface appears at precisely that moment and changes the rhythm of work: fewer questions, more time, better recovery. Firms hire people with similar technical labels, but the productivity association appears where those workers build knowledge the organization can reuse.
The scarce input in each case is neither the model nor information alone. It is translation capital: enough context to ask the right question, enough judgment to evaluate the answer, enough institutional permission to use it, and a workflow capable of making the result stick.
Economists usually expect a cheap general-purpose input to diffuse through prices. But a model response is not yet an economic good. It becomes one only when somebody trusts it, checks it, and has the right to act. Those rights are unevenly distributed. Junior workers may use AI but lack authority to alter the process. Households may receive excellent advice but lack the confidence or liquidity to follow it. Managers may understand a better allocation and still hoard the employee because the bonus rewards local performance. The bottleneck sits inside the institution.
This changes the strategic question for both governments and firms. “How many people have access?” is rapidly becoming the wrong metric. The better denominator is the number of relevant decisions, errors, claims, or handoffs. The numerator is how many become verified, completed actions. Measured that way, an enterprise with universal licenses can have near-zero adoption where it matters. A public agency with a modest domain-specific assistant can create more capacity than a ministry that merely offers a frontier chatbot.
It also changes the inequality question. If access prices approach zero, advantage does not disappear. It moves downstream: into framing, verification, confidence, execution, and control over the interface. That is why equal access can raise average welfare and still widen a gap. The advantaged do not necessarily own the better model. They own more of the conversion chain.
The policy implication is not to automate every action by default. That would substitute a new concentration of agency for the old one. It is to make the conversion chain observable and contestable: disclose where recommendations come from, preserve escalation, measure completed outcomes across groups, and separate permission to advise from permission to execute.
The organizations that understand this will stop buying “AI” as a noun. They will redesign verbs: diagnose, compare and approve; appeal, teach and verify. That is where productivity will appear first. It is also where power will move first.
09B
Implications
What changes when the interface becomes an institution?
Growth and productivity
The August papers strengthen a J-curve interpretation without proving an aggregate boom. Firms may incur a measured cost while building evaluation routines, data assets, and redesigned processes. Productivity statistics should therefore separate current AI use from organization-capital formation and should look for heterogeneity across firms rather than only an economy-wide mean.
Labor, wages, and human capital
Hiring is a leading indicator that headcount and wages can miss. Researchers and firms should track entry rates, task reassignment, promotion pipelines, and skill formation by cohort. If junior jobs are training technologies, subsidizing or redesigning apprenticeships may become more relevant than protecting every incumbent task.
Firms, investment, and business strategy
Buying model access is the easy investment. The scarce investments are process ownership, validation, escalation, data rights, and incentives for managers to share talent and codify knowledge. A useful business metric is not seats licensed but verified decisions or completed actions per exposed workflow, alongside error and override rates.
Development and distribution
The World Bank estimates that tasks in 4.5% of jobs in developing economies face automation potential, compared with 14.2% in high-income economies; augmentation potential is closer, at 16.2% versus 18.7%. Lower displacement exposure is not automatically an advantage if it reflects poor infrastructure and low adoption. Distribution policy must address electricity, connectivity, local language, identity systems, and trusted execution—not only provide access to a model.
Competition and market structure
If the interface determines conversion, control over defaults, workflow data, and execution rights becomes a source of rent. Model competition alone may not discipline an incumbent that owns the distribution surface or the proprietary feedback loop. Interoperability and data portability could matter as much as benchmark rivalry, although August offers mechanism evidence rather than a market-power estimate.
Fiscal capacity and public administration
The tax experiment shows that AI can shift a real administrative action at low marginal cost. Public agencies could use assistants to make rights usable, not merely publish information. But a system that disproportionately helps citizens already equipped to engage can widen effective access to the state. Distributional audits should therefore measure completed claims by group, not chatbot availability.
Regulation and staged deployment
Gans's staged-access model treats release timing and liability as complements: defensive effort rises as release approaches, delay has diminishing value, and liability alone need not implement the socially preferred timing. The practical implication is to evaluate deployment as a sequence of permissions—advice, drafting, recommendation, execution—each with different verification and liability requirements. This is a theoretical result, not evidence that one staging rule fits every sector.
10
Serious contributions screened
The complete 36-record research ledger
Scores guide editorial attention; they are not cardinal estimates of scientific quality.
Open source appendix
Score
Paper
Status
Topic
Use
Source
87
AI Financial Advice: Supply, Demand, and Life Cycle ImplicationsTaha Choukhmane; Tim de Silva; Weidong Lin; Matthew Akuzawa
Prompt experiment plus simulation; working paperarXiv; NBER WP 35574
AI, household finance
Yes
90
World Development Report 2026: The Promise of Artificial IntelligenceWorld Bank
Institutional report and task modelingWorld Bank
AI, development
Yes
86
Understanding Firms’ AI Efforts and Their Economic ImpactTania Babina
Corrected proof; review articleReview of Corporate Finance Studies
AI measurement, firms
Yes
89
Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial IntelligenceErik Brynjolfsson; Bharat Chandar; Ruyu Chen
Descriptive administrative-data working paperStanford Digital Economy Lab
AI, labor
Yes
90
Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice ExperimentPhilip Oreopoulos; Michael Liut; Alp Sungu; Nina Low
Randomized experiment; working paperEdWorkingPaper 26-1552; NBER WP 35621
AI, education
Yes
94
Taxpayer Behavior in the Age of AI: A Field Experiment on Property Tax AppealsJustin E. Holz; Ricardo Perez-Truglia; Andrew Simon; Alejandro Zentner
Randomized field experimentNBER WP 35632
AI, public economics
Yes
93
What Work Does Generative AI Do?Alexander Bick; Adam Blandin; David J. Deming; Tyler R. Schumacher
National survey and task linkage; working paperNBER WP 35677; CEPR DP 21889
AI adoption, measurement
Yes
92
One Click Away: AI Tutoring with Khanmigo in a Two-Year School ExperimentPhilip Oreopoulos; Nina Low
Virtual Tutoring with Computer-Assisted Learning: An Experiment in Take-Up and LearningPhilip Oreopoulos; Ruochong Dong; Nina Low
Two-year randomized offerNBER WP 35622
Education, take-up
Yes
88
Canaries in the Gold Mine: Early Productivity Gains from Artificial Intelligence Creating Organization CapitalTania Babina; Alex X. He; Renhao Jiang
Observational firm panel; working paperNBER WP 35684; CEPR DP 21894
AI, productivity
Yes
85
AI Agents and Prompt Engineering in Econometric CodingSebastián Galiani; Ariel López; Raul A. Sosa
Repeated coding benchmark; working paperNBER WP 35588
AI, research production
No
83
Staged Access and Liability for Dual-Use Artificial IntelligenceJoshua S. Gans
TheoryNBER WP 35586
AI governance
Yes
82
The Global Transition – The Impact of Demographics and AI on Economic PowerSeth G. Benzell; Laurence J. Kotlikoff; Victor Yifan Ye
Calibrated global scenariosNBER WP 35618
AI, demography, growth
Yes
79
When the Middle Class Undermines Progress: The Political Economy of AI RegulationAnna Denisenko; Konstantin Sonin
TheoryBecker Friedman Institute; CEPR DP 21888
AI, political economy
No
80
Artificial Intelligence and Political AdviceGeorgy Egorov; Konstantin Sonin
TheoryBFI; NBER WP 35689
AI, information, politics
No
80
Endogenous Task Bundling, Skills and AutomationJoshua S. Gans
Theory; revised working paperNBER WP 35211
Automation, tasks
No
94
Clause and Effect: Theory and Field Experimental Evidence on Noncompete ClausesBo Cowgill; Brandon Freiberg; Evan Starr
Accepted article; field experimentQuarterly Journal of Economics
Labor mobility
Yes
93
Sticky Wage Norms and the Real Wage Cost of Unexpected InflationErik Hurst; Christina Patterson; Nela Thomas Richardson; Ye Liv Wang
Administrative payroll study; working paperNBER WP 35624; BFI
Inflation, wages
Yes
93
The Employment Effects of a Guaranteed Income: Experimental Evidence from Two U.S. StatesEva Vivalt; Elizabeth Rhodes; Alexander Bartik; David Broockman; Patrick Krause; Sarah Miller
Accepted article; three-year RCTQuarterly Journal of Economics
Additionality and Asymmetric Information in Environmental Markets: Evidence from Conservation AuctionsKarl M. Aspelund; Anna Russo
Peer-reviewed empirical and mechanism-design studyAmerican Economic Review 116(8)
Environment, auctions
No
89
Uncertainty and Change: Survey Evidence of Firms’ Subjective BeliefsRüdiger Bachmann; Kai Carstensen; Stefan Lautenbacher; Manuel Menkhoff; Martin Schneider
Do Children Derail Women from the CEO Track?Santiago Campos-Rodríguez; David Neumark
Longitudinal administrative studyNBER WP 35616
Gender, careers
No
88
Declining Occupations and Career Outcomes in the United StatesErling Barth; Maria Forthun Hoen; Sari Pekkala Kerr; William R. Kerr
Linked census-administrative panelNBER WP 35614
Structural change, careers
No
Verification record
Research-period and source audit
Primary period: 1–31 August 2026. Cutoff: 1 September 2026.
Journal records were checked against the AEA and Oxford Academic article pages; working-paper titles, numbers, authors, and dates were checked against NBER, CEPR, BFI, EdWorkingPapers, Stanford, arXiv, or the World Bank.
Numerical claims in the narrative were retained only when traceable to an abstract, official release, paper text, table, figure, or author manuscript. No media report is used as the evidentiary source for a result.
Causal language is reserved for randomized or defensibly quasi-experimental variation. Administrative patterns are labeled descriptive; firm associations observational; model outputs structural or simulated; institutional task shares modeled potential rather than forecasts.
The appendix contains 36 screened contributions. Twelve form the main paper shortlist; additional items enter the focal synthesis where their mechanism or measurement contribution is material.
Date and status exceptions
Item
Why it appears in an August Radar
Treatment in this report
What Work Does Generative AI Do?
Earlier public visibility; formal NBER and CEPR series entry in August, with CEPR dated 31 August.
Included; status disclosed.
Canaries in the Coal Mine?
Originally circulated in 2025; materially revised 12 August with payroll data through June 2026.
Included as a revision, not misdated as a new paper.
Understanding Firms’ AI Efforts and Their Economic Impact
Published 27 July; corrected and typeset 7 August.
Appendix and measurement framework only.
Endogenous Task Bundling, Skills and Automation
Earlier NBER circulation; substantive August revision.
Appendix only.
Claim-level checks
Units and baselines: Treatment-control differences, percentage levels, percentage-point changes, percentile ranks, standard deviations, counts, and simulated wealth differences remain explicitly distinguished.
Denominators: The tax, Khanmigo, NUMI, task-adoption, labor, wage, and World Bank numbers are not pooled. The chart specifications prohibit funnels, Sankeys, or common axes where the population or unit changes.
Uncertainty: The NUMI delayed-learning estimate is labeled marginally significant; the virtual-tutoring test estimate is reported with a confidence interval spanning zero in the appendix; the 19% employment result is labeled descriptive and sensitivity-dependent.
Scenarios: The financial-advice wealth gaps and global-transition results are identified as simulations or calibrated scenarios, not realized wealth or forecasts.
Unverified material: None retained. Any claim that could not be tied to a stable primary record was excluded rather than marked for publication.
Radar Signal Score rubric
Dimension
Weight
Scientific rigor
25%
Originality
20%
Economic significance
20%
Potential long-term importance
15%
Policy and business relevance
10%
Surprise or challenge to conventional wisdom
10%
The score supports selection; it does not imply that 92 is statistically distinguishable from 89. Randomized designs, transparent administrative data, direct behavioral measurement, and consequential outcomes receive credit. Weak identification, narrow external validity, preliminary status, model dependence, or scenario assumptions receive explicit penalties.
Publication QA
Run-specific placeholders: cleared.
Calendar window and cutoff: stated.
Publication, first-circulation, and revision dates: distinguished where material.
Evidence types and causal status: labeled.
Quantitative claims, units, denominators, and baselines: checked against primary records.
Markdown headings, tables, links, and section order: checked.
Duplicate papers and duplicate findings: removed or cross-referenced.