July 2026 · Economics of AI

Implementation Became the Mechanism

AI productivity is a conversion problem, not a capability statistic.

July moved the economic frontier from model capability to conversion: organizational design, verification, market structure, and institutional scarcity determine whether new technology becomes measured output.

2026-07-01 to 2026-07-31 12 ranked papers Cutoff 2026-08-12
01

Executive signal

Three corrections define the month

July's research did not produce one dominant answer. It produced three unusually useful corrections.

The first concerns growth. Two ambitious papers challenge comfortable extrapolations from the past, but in opposite directions. Jones, López-Salido, and Philippon argue that US total factor productivity has historically advanced by roughly constant additions to its level, not by a constant exponential rate. That turns the familiar secular slowdown into something closer to an arithmetic implication. De Souza and coauthors, by contrast, estimate that US public R&D generates very large and slow-moving productivity spillovers abroad. Read together, the papers suggest that trend growth may be less automatic than standard models assume and more dependent on deliberately produced, internationally diffusing knowledge than national accounts reveal.

The second correction concerns scarcity. Acemoglu, Autor, Beirne, and Scott find that lower birth rates have been followed by faster growth in output per working-age adult, without a detectable loss of aggregate output. Their proposed mechanism is not a free demographic lunch. Scarcity of younger workers changes the direction of innovation and induces labor-saving investment. This is a more dynamic account of aging than merely subtracting workers from a fixed production function.

The third—and the month's clearest AI signal—concerns the unit of analysis. AI capability, tool adoption, task speed, worker output, firm value added, and aggregate productivity are not interchangeable variables. Pakistan's randomized JudgeGPT rollout raises completed cases by 6.3% only when access is paired with task-specific training. A hospital experiment shows that developer guardrails can redirect real prescriptions and tests without establishing whether health improves. A student experiment finds a 0.27-standard-deviation knowledge gain, but durable benefits are concentrated among users who ask AI to explain rather than to do the work. A Federal Reserve measurement note still sees a buildout phase, not broad transformation. And a competing aggregate interpretation attributes essentially all US TFP growth since early 2024 to higher utilization rather than better underlying efficiency.

The intellectual change is therefore subtle but consequential: the frontier has moved from asking what models can do to asking which production systems can convert those capabilities into verified economic output. The scarce input is increasingly not intelligence in the abstract. It is organizational design.

Across the broader radar, the same logic reappears. A policy label or a newly installed tool is not itself a treatment. Union incidence changes with market power; protected status without strict rules produces little visible environmental change; communication technology requires managerial incentives; and an industrial big push can become long-run lock-in. July was a month in which implementation stopped being a footnote and became the mechanism.

02

Economics Radar

A scored field, not a news feed

Scores are editorial decision aids. Evidence labels remain more important than small rank differences.

94ModerateAccelerating

Public R&D

Foreign TFP rises about 1% after a 1%-of-capital US R&D shock

National appraisal misses large global spillovers

93ModerateWeakening

Growth

Additive TFP dominates exponential fit in long historical data

Constant percentage-growth assumptions need justification

93StrongReversing

Industrial clusters

A large early productivity gain reverses into a long-run disadvantage

Anchor plants can crowd out renewal rather than seed it

92ModerateReversing

Demography

Lower birth rates predict higher output per worker without lower aggregate GDP

Scarcity can redirect innovation

92StrongAccelerating

AI implementation

JudgeGPT raises output only with targeted training

Organizational capital is part of the technology

92StrongAccelerating

AI guardrails

Patient-side AI lowers prescriptions but raises diagnostic testing

Developer objectives can propagate into real expert decisions

92StrongReversing

Replication

Corrected inventor-cluster causal estimates are insignificant

Place-based policy loses a key calibration

91ModerateAccelerating

Inequality

Holding companies conceal roughly half of top-0.1% income

Legal realization distorts progressivity and inequality

90ModerateUncertain

Inflation

Apparent anchoring can arise without target credibility

Stability may unravel after repeated shocks

90EarlyAccelerating

AI markets

Usage-based market exposure differs from technical task exposure

Incidence is endogenous to use and market structure

88ModerateAccelerating

AI learning

Augmentation builds more durable knowledge than automation

Human-capital effects depend on use mode

88EarlyUncertain

AI investment

Contest model implies 1.5× efficient investment in baseline

Private races can create financial fragility

87ModerateReversing

Evidence synthesis

Selection correction can reduce mean effects to 12–21% of naïve averages

Literature means can greatly overstate treatment effects

03

Frontier observatory

The shape of the evidence

The charts describe this edition, not an imagined economy: signal scores, evidence labels, paper rankings, and the screened research ledger.

13 signals12 ranked papers70 screened records
Signal topologyScore / editorial sequence
1009080
Moderate evidence94

Public R&D

Foreign TFP rises about 1% after a 1%-of-capital US R&D shock

National appraisal misses large global spillovers
Evidence composition13 signals
Strong
4
Moderate
7
Early
2
Source terrainLargest topic clusters
  1. Economics of AI20 records · mean score 86
  2. Macro & growth18 records · mean score 88
  3. Human capital & health10 records · mean score 87
  4. Labor & distribution6 records · mean score 90
  5. Methods, institutions & behavior5 records · mean score 88
  6. Firms & industrial change4 records · mean score 88
Paper skylineRadar score by rank
04

The papers that matter

Twelve contributions worth carrying forward

Each record separates status, method, result, relevance, and limitation.

  1. 0194
    Moderate — serious working-paper or institutional evidenceNBER Working Paper 35517, issued 27 July 2026; DOI 10.3386/w35517.

    The Effects of U.S. Public R&D on Global Growth

    Gustavo De Souza (Federal Reserve Bank of Chicago), Andrew J. Fieldhouse (Texas A&M Mays Business School), Karel Mertens (Federal Reserve Bank of Dallas), Ishan B. Nath (Harvard Kennedy School and NBER), and Valerie A. Ramey (Stanford, Hoover Institution, and NBER).

    A shock equal to 1% of the US federal R&D capital stock raises foreign TFP by approximately 1% after 12 years. The peak response is about 1.8% in non-OECD economies and 0.7% in OECD economies.

    Method, implication, and boundary
    Question
    How much does federally funded US R&D raise productivity outside the United States?
    Method / data
    Narrative identification of exogenous federal R&D appropriations shocks; local projections for 69 foreign economies, 1980–2019.
    Why it matters
    It supplies quasi-experimental evidence for an international knowledge externality large enough to change optimal R&D policy and burden sharing.
    Limitations
    Long-horizon narrative instruments require credible exclusion; cross-country TFP is noisy; the diffusion channels are inferred rather than individually identified.
    Primary record
  2. 0293
    Moderate — serious working-paper or institutional evidenceNBER Working Paper 35415, issued 6 July 2026; DOI 10.3386/w35415. Paper dated June 2026.

    Is Growth Additive?

    Callum J. Jones and David López-Salido (Federal Reserve Board); Thomas Philippon (New York University and NBER).

    Additive TFP describes approximately 90 years of US evidence better than the geometric benchmark. Postwar TFP adds roughly 2.5% of its 1947 level annually, implying a declining percentage growth rate as the level rises.

    Method, implication, and boundary
    Question
    Does productivity grow by a constant percentage or by a roughly constant addition to its level?
    Method / data
    US TFP from 1890–2022, professional forecasts, international data for 23 countries plus the euro area, out-of-sample comparison, and Bayesian model selection.
    Why it matters
    Many growth models, forecasts, fiscal projections, and AI scenarios build in exponential productivity as a default. This paper turns that default into a testable—and rejected—hypothesis.
    Limitations
    TFP is a residual; structural breaks and historical measurement matter; a reduced-form law of motion does not explain why growth is additive or rule out an AI-induced regime change.
    Primary record
  3. 0392
    Moderate — serious working-paper or institutional evidenceNBER Working Paper 35401, issued July 2026; DOI 10.3386/w35401. Manuscript dated 24 June; SSRN revision 20 July.

    Baby Busts and Growth Booms: Demographic Change and the Macroeconomy

    Daron Acemoglu, David Autor, and Keelan Beirne (MIT); Andrew Scott (London Business School and Ellison Institute of Technology).

    Lower birth rates predict higher GDP per working-age adult and wages, with no detectable fall in aggregate GDP or earnings. The evidence points to labor-saving technological change.

    Method, implication, and boundary
    Question
    Do lower birth rates and an older workforce reduce long-run growth?
    Method / data
    Seven decades of country data, 722 US commuting zones, industry outcomes, patent classifications, and cross-country WWII mortality variation.
    Why it matters
    It challenges mechanically pessimistic aging scenarios and restores directed innovation to demographic analysis.
    Limitations
    Long-run correlations, migration, and measurement make causal interpretation demanding; patent and high-tech proxies do not fully identify adoption or welfare.
    Primary record
  4. 0492
    Strong — peer-reviewed or accepted researchAccepted/online-first Quarterly Journal of Economics, published 1 July 2026; DOI 10.1093/qje/qjag034.

    Who Pays for Unions?

    Samuel Dodini (Federal Reserve Bank of Dallas), Anna Stansbury (MIT Sloan), Alexander Willén (Norwegian School of Economics).

    In the average private firm, higher unionization raises labor cost but reduces employment, output, and profit; labor's share does not rise. In manufacturing and less competitive markets, wages, employment, and output rise while markdowns fall.

    Method, implication, and boundary
    Question
    Who bears the incidence when union density rises?
    Method / data
    Norwegian administrative employer–employee data and tax-deductibility reforms that shift the price of union membership.
    Why it matters
    It joins labor and industrial organization: union effects depend on where rents exist and how product markets transmit costs.
    Limitations
    Norwegian institutions and reform compliers bound external validity; competition is measured rather than randomized.
    Primary record
  5. 0592
    Strong — peer-reviewed or accepted researchPeer-reviewed comment, American Economic Review 116(7), July 2026, 2754–2763; DOI 10.1257/aer.20231415.

    The Effect of High-Tech Clusters on the Productivity of Top Inventors: Comment

    Michael Wiebe (independent economist).

    The original 0.0676 baseline elasticity is not supported by corrected mover-event-study and IV estimates, which are statistically insignificant. Errors include an omitted, correctly interacted event-time-zero treatment and an IV differenced across cities after unsorted data.

    Method, implication, and boundary
    Question
    Does the influential Moretti estimate of cluster-size effects on top inventors survive code and specification correction?
    Method / data
    Direct replication and correction of mover-event-study and IV code.
    Why it matters
    Agglomeration estimates influence place-based innovation policy. A transparent correction is a public good even when it yields a null.
    Limitations
    The corrected mover sample is roughly 3,000 inventors versus around 118,000 in baseline OLS. Insignificance in the smaller causal samples does not prove zero agglomeration spillovers.
    Primary record
  6. 0691
    Moderate — serious working-paper or institutional evidenceSubstantially revised NBER Working Paper 35504 and New York Fed Staff Report 1198, 20 July 2026; DOI 10.59576/sr.1198. An earlier version circulated in January.

    Bank Runs With and Without Bank Failure

    Sergio A. Correia (Federal Reserve Bank of Richmond), Stephan Luck (Federal Reserve Bank of New York), and Emil Verner (MIT Sloan and NBER).

    Runs hit strong and weak banks, but weak pre-run fundamentals predict failure and coincide with larger local lending and real-activity contractions. Bank fundamentals alone predict runs with an AUC of about 0.69; adding aggregate and local information raises it to 0.80.

    Method, implication, and boundary
    Question
    What distinguishes runs that reveal weakness from runs that destroy otherwise viable banks?
    Method / data
    LLM-assisted extraction and human validation of US newspaper archives, producing 3,984 bank-run events, 1863–1934.
    Why it matters
    It reconciles panic and fundamentals accounts: a run can be a coordination event while its social cost remains state-dependent.
    Limitations
    Pre-FDIC banking differs from today's regime; newspaper coverage and generated classifications are selective despite validation; the paper cannot fully separate the effect of a run from the underlying local real shock.
    Primary record
  7. 0791
    Moderate — serious working-paper or institutional evidenceNBER Working Paper 35534, issued 28 July 2026; DOI 10.3386/w35534.

    Personal Holding Companies, Tax Progressivity, and Inequality

    Marius A. K. Ring (University of Texas at Austin), David Seim (Stockholm University), and Gabriel Zucman (Paris School of Economics and UC Berkeley).

    Around half of top-0.1% income is retained in PHCs; the highest-income owners distribute only 15–20% over 20 years. Effective total tax rates fall from around 50% in the upper middle to 15–20% at the top of wealth.

    Method, implication, and boundary
    Question
    How much top income is retained inside personal holding companies, and what does that do to measured inequality and tax progressivity?
    Method / data
    Twenty years of linked Swedish and Norwegian owner–firm administrative data; event studies around operating-company value-added shocks.
    Why it matters
    The legal perimeter of the tax unit can make a progressive schedule economically regressive at the top and distort international comparisons.
    Limitations
    Private-firm valuation and imputed accrual income introduce uncertainty; the Nordic systems are distinctive.
    Primary record
  8. 0892
    Strong within the study setting; external validity remains boundedCEPR Discussion Paper 21783, 23 July 2026; revised manuscript dated 9 August; SSRN DOI 10.2139/ssrn.7119447.

    Courts of Tomorrow: Evidence from a Nationwide Rollout of Generative AI

    Sultan Mehmood (New Economic School); Christoph Goessmann and Elliott Ash (ETH Zurich; current manuscript affiliations).

    At median district exposure, AI plus targeted training produces 1,848 additional resolved cases per district-year, 6.3% above the mean. Appeals slightly decline; measured writing quality does not deteriorate. Targeted instruction increases use at least fourfold relative to generic training and directs it toward bounded, verifiable tasks.

    Method, implication, and boundary
    Question
    Can generative AI raise output in a high-stakes public organization, and does task-specific training matter?
    Method / data
    Preregistered randomized rollout to 1,559 judges in 118 Pakistani courts: JudgeGPT with targeted training, JudgeGPT with generic training, or generic training without access; administrative case records, appeals, opinions, and chat logs.
    Why it matters
    This is the month's strongest causal evidence that organizational complementarity converts model capability into institutional capacity.
    Limitations
    Treatment exposure is aggregated at district level; quality measures cannot capture every legal error; Pakistan's congested, low-staffed courts are not a representative organization.
    Primary record
  9. 0990
    Moderate — serious working-paper or institutional evidenceNBER Working Paper 35451, July 2026; DOI 10.3386/w35451.

    AI Premium

    Nicola Borri (LUISS), Yukun Liu (University of Rochester, Simon Business School), and Aleh Tsyvinski (Yale and Cowles Foundation).

    A value-weighted high-minus-low AI-beta portfolio earns 64.1 basis points per week (t=2.84), about 56 after standard factor controls (t=2.41). Interaction and communication skills load positively; analytical and information-use skills load negatively. Established engineering exposure measures explain under 2% of cross-firm beta variation.

    Method, implication, and boundary
    Question
    Can realized AI usage produce a tradable economy-wide factor and reveal which firms and skills markets perceive as exposed?
    Method / data
    Licensed OpenRouter data covering 380 trillion tokens, more than 400 models, millions of accounts, January 2024–April 2026; high-frequency AI factors, firm betas, portfolio sorts, event studies.
    Why it matters
    It measures exposure from use and prices rather than occupational possibility scores, revealing a different map of the AI economy.
    Limitations
    Two years is short for an asset-pricing premium; the factor is a generated regressor from one intermediary; risk, mispricing, and cash-flow channels remain unresolved.
    Primary record
  10. 1088
    Moderate — serious working-paper or institutional evidenceIZA Discussion Paper 18792, July 2026.

    Experimental Evidence on the Learning Impact of Generative AI

    Zara Contractor and Germán Reyes, Middlebury College; Reyes also IZA@LISER.

    AI access raises immediate knowledge scores 6.7 percentage points from a 56.3% control mean, equal to 0.27 standard deviations. One week later, the gain is 5.1 points—76% of the immediate effect. Augmentation users retain gains; automation users' assisted essay gains disappear.

    Method, implication, and boundary
    Question
    Does off-the-shelf generative AI improve durable learning, or merely assisted performance?
    Method / data
    Proctored RCT with 211 undergraduates and 204 one-week returnees; immediate and delayed unaided tests and essays.
    Why it matters
    It separates output today from human capital tomorrow and makes use mode—not access—the relevant treatment margin.
    Limitations
    Small, selective college sample; three topics; one hour of treatment; augmentation/automation classification is endogenous.
    Primary record
  11. 1193
    Strong — peer-reviewed or accepted researchOnline-first/corrected proof, Review of Economic Studies, 20 July 2026; DOI 10.1093/restud/rdag078.

    Industrial Clusters in the Long Run: Evidence from Million-Rouble Plants in China

    Stephan Heblich (University of Toronto and NBER), Marlon Seror (UQAM), Hao Xu (China Construction Bank), and Yanos Zylberberg (University of Bristol and CEPR).

    Host counties were about 80% more productive than controls in 1982 but 20% less productive by 2010. Early industrialization raised non-agricultural household registration by roughly 30 percentage points, yet later entrants were low-productivity, innovation was limited, and markups were high.

    Method, implication, and boundary
    Question
    Does a state-created industrial cluster generate durable local productivity spillovers, or can early specialization become long-run lock-in?
    Method / data
    The 1950s Soviet-assisted “Million-Rouble Plants” program: 149 large plants across 98 Chinese counties. The design compares suitable locations and uses historical airbase proximity, which affected geopolitical vulnerability and siting, to isolate plant placement.
    Why it matters
    It separates a successful factory from a successful local ecosystem. A genuine big push can raise short-run productivity while crowding out entrepreneurship and adaptation over the longer run.
    Limitations
    The planned-to-market transition in China is unusual; the instrument and suitable-location comparison require the stated exclusion assumptions; county estimates include local-equilibrium effects.
    Primary record
  12. 1292
    Strong within the study setting; external validity remains boundedSIEPR Working Paper and arXiv:2607.08706, first posted 9 July 2026; manuscript dated 10 July. Preregistered as AEARCTR-0015851.

    Directional AI Advice: Experimental Evidence from Healthcare

    Yuyu Chen (Peking University), Hongbin Li and Lingsheng Meng (Stanford University), and Xinyao Qiu and Qingxu Yang (University of Hong Kong).

    Randomized access lowers the probability of any prescription by 4.6 percentage points from an 87% baseline and raises diagnostic testing by 2.7 points from a 23% baseline. Total spending is unchanged. The chatbot cautions against medication but usually recommends diagnostic testing, and clinical choices move in the same direction.

    Method, implication, and boundary
    Question
    Do the directional guardrails embedded in a consumer AI system propagate into real decisions between patients and physicians?
    Method / data
    Two-layer field randomization at a large Chinese public hospital: 179 physicians and 17,676 completed outpatient visits, linked to prescriptions, tests, spending, revisits, surveys, and 1,192 chatbot turns. Of 5,828 visits offered access, 17.1% used the chatbot.
    Why it matters
    It shows that model-governance choices can alter real expert–client decisions even when the AI is used by the client rather than the professional. The relevant “AI treatment” includes developer guardrails and organizational authority.
    Limitations
    One hospital and one national context; online-booking patients only; low take-up makes the IV a local effect; no clinical health endpoint; the experiment cannot sign welfare because fewer drugs and more tests may each be beneficial or wasteful.
    Primary record
05

Focus · Economics of AI

The conversion problem

Capability, adoption, task performance, organizational output, and aggregate productivity are different variables.

  1. 01Capability

    Strong evidence of technical progress; weak mapping to whole jobs

  2. 02Adoption

    Strong on direction, moderate on comparable levels

  3. 03Task productivity

    Strong in bounded tasks; heterogeneous beyond them

  4. 04Worker productivity

    Moderate-to-strong in selected occupations

  5. 05Expert–client decisions

    Strong on decision direction; welfare remains unidentified

  6. 06Firm productivity

    Moderate; causal value-added evidence remains scarce

  7. 07Aggregate productivity

    Early and disputed

  8. 08Distribution

    Early; realized wage and incidence effects remain thin

Where the literature stood

By June 2026, five propositions had become reasonably defensible.

First, frontier capability was improving rapidly, especially on bounded language, coding, and analytical tasks. Second, access to generative AI often raised task-level speed or quality in experiments. Third, adoption was broad but shallow: many workers had tried the tools, while only a small share of total hours and complete workflows used them. Fourth, effects were heterogeneous across the “jagged frontier”: AI could help weaker performers on well-defined tasks and harm performance where contextual judgment or verification exceeded user skill. Fifth, aggregate productivity and employment data did not yet contain an unambiguous AI break.

This left a translation puzzle. The literature could show impressive capability and credible task gains but offered much thinner evidence on worker output over time, firm value added, organizational redesign, and economy-wide TFP. July's contribution is not to close that gap. It is to identify the conversion mechanisms inside it.

What changed this month?

1. Implementation became an estimand

The JudgeGPT experiment does something most AI evaluations do not: it randomizes not only access but the complement that tells users where and how to deploy it. Targeted training increases usage at least fourfold relative to generic technology training, moves prompts away from costly-to-verify open legal questions, and produces a 6.3% output increase without a measured quality loss. Access alone is not the relevant treatment.

This finding aligns with the non-AI garment-factory RCT in the broad radar. A communication technology creates no value until HR incentives change. The common mechanism is organizational: the effective productivity of a tool depends on whether someone owns adoption, workers know the tool's comparative advantage, and verification is assigned to tasks where it is economical.

2. Guardrails became an economic input

Chen et al. randomize patient access to a medical chatbot before outpatient visits. Only 17.1% of those offered access use it, yet the intent-to-treat effects move real care: prescriptions fall 4.6 percentage points from an 87% baseline, while diagnostic testing rises 2.7 points from a 23% baseline. Total spending is unchanged.

The conversation logs reveal the mechanism. The chatbot attaches cautions to 90.8% of traditional-Chinese-medicine mentions and 87.6% of antibiotic mentions, but gives a clean recommendation in 94.5% of diagnostic-test discussions. The system's defensive, liability-sensitive design propagates into clinical choices, especially when physicians are receptive to patient input. This is not evidence that AI improved health: the paper has no clinical outcome with which to value fewer drugs against more tests. It is strong evidence that the developer's implicit loss function can become an input into another institution's production process.

3. Durable human capital depends on augmentation, not mere assistance

Contractor and Reyes distinguish assisted output from later unaided performance. Their 0.27-standard-deviation immediate knowledge gain retains 76% after one week. But use logs reveal the boundary: students who ask for explanations preserve benefits; students who outsource text see short-run essay gains vanish when the tool is removed. AI can be a tutor or a substitute for practice. “Used AI” is too coarse a variable to tell which.

4. Realized usage produces a different map from technical exposure

AI Premium uses 380 trillion tokens of actual consumption to infer an AI factor and firm-level market betas. Technical exposure measures explain under 2% of their cross-firm variation. Market exposure is more positive for interaction and communication content and more negative for analytical and information-use skills. That is surprising only if capability maps are mistaken for equilibrium incidence. Prices reflect adoption, complementarity, competition, bargaining, and expected rents—not merely whether an LLM can perform a task.

5. More compute is not monotonically more useful

Havránek and Irsova's negative experiment matters beyond academic feedback. Authors prefer a single frontier-model pass to two multi-agent debate systems, although one uses around 30 times the tokens. AI judges almost always rank the real human referee report last. The result is narrow but sharp: inference-time complexity can lower user value, and model-based evaluation can reverse the human outcome criterion.

6. The buildout itself may contain a contest externality

The BIS AI Investment Race calibrates a dynamic winner-take-most contest among hyperscaler–lab coalitions. Its baseline produces investment at 1.51 times the efficient level, a 50% bust probability, and expected destruction of $189 billion per year. These are model outputs, not forecasts. Still, the mechanism is important: private investment can be individually rational and collectively excessive when firms race for a few dominant positions, finance specialized assets with debt, and own circular stakes in one another.

Contradictions

The macro evidence is not yet internally consistent—and should not be made to look so.

Scott Davis reports that the three most AI-exposed US sectors grew productivity by 3.7% annualized since early 2024 versus 1.7% elsewhere. They account for 16% of hours but 40% of US productivity gains. Across European countries, greater Claude use is associated with a stronger exposure–productivity slope. This is suggestive of realized AI gains, but explicitly correlational.

Boyle, Fernald, and Li offer a competing decomposition. US output per hour grew about 2.5% annually from 2023 to 2026Q1, one point above 2005–19. Measured TFP contributed 0.8 point to the acceleration, yet since early 2024 their utilization estimate accounts for essentially all TFP growth, leaving little utilization-adjusted improvement. In this reading, AI may have raised demand and uncertainty, causing firms to work existing inputs harder—not yet smarter.

Both can be true. AI-exposed sectors may be improving while aggregate TFP is dominated by cyclical intensity; or exposure may proxy for pre-existing sectoral advantages. The decisive evidence will be sustained, utilization-adjusted productivity divergence within industries after comparable adoption, ideally linked to firm-level workflows and value added.

Measurement problems

1. Capability is not economically weighted output

Benchmarks assign equal or arbitrary weight to tasks. GDP weights outputs by market value; welfare also includes time, quality, variety, consumer surplus, and nonmarket services. Saturating an exam does not reveal which production bottleneck has been relaxed.

2. Adoption is usually binary, intensity is continuous

“Uses AI” can mean one query per month or a redesigned production process. The Fed note finds uptake rising with firm size but emphasizes shallow intensity. Surveys need task shares, workflow penetration, model costs, verification time, and complementary investment.

3. Quality adjustment is especially hard in services

Courts can count cases, but justice quality is multidimensional. Education can count test answers, but durable reasoning differs from polished text. In many services, nominal revenue is deflated with imperfect prices, so free or cheaper AI-enhanced quality can disappear from measured real output.

4. AI capital straddles accounting categories

Data centers and computers enter investment; imported equipment subtracts through net exports; software is partly capitalized; data, model fine-tuning, training, and workflow redesign are incompletely measured intangibles. The Fed estimates that AI-related components contributed meaningfully to GDP growth from 2025 to 2026Q1, while imports offset much of gross investment in some quarters.

5. Exposure measures answer different questions

Engineering scores ask what AI could perform. Usage logs ask what people currently delegate. Stock betas ask which cash flows or discount rates covary with AI demand. These measures should disagree. Treating one as ground truth creates false contradictions.

6. Generated measurement can contaminate downstream inference

July's bank-run archive shows the upside of LLM extraction. But AI Premium, AI-judged paper feedback, text-based court-quality scores, and generated covariates all place model error inside the estimator. Validation must be designed around the target coefficient, not generic classification accuracy.

7. Utilization is not welfare

Fewer prescriptions, more diagnostic tests, more resolved cases, or more books are intermediate outcomes. Their value depends on clinical health, legal accuracy, consumer use, and the counterfactual cost of resources. AI research will overstate progress if it counts directional behavior as improved welfare without measuring the endpoint that gives the behavior economic value.

06

Numbers worth remembering

Ten quantities with their caveats attached

  1. 6.3%

    additional court cases resolved at median district exposure to JudgeGPT with targeted training, not AI access in isolation.

  2. 0.27 SD

    immediate knowledge gain from AI access in the Middlebury experiment; 76% persists one week later.

  3. 1% after 12 years

    estimated foreign TFP response to a US R&D appropriations shock equal to 1% of federal R&D capital.

  4. 0.993

    minimum reported posterior probability on the additive TFP model by 2022 across prior choices.

  5. 26.8%

    higher GDP per working-age adult associated with a one-point lower birth rate across countries over 1970–2020.

  6. 64.1 basis points per week

    value-weighted high-minus-low AI-beta return in AI Premium; the short sample makes this a signal, not a settled premium.

  7. 1.51×

    BIS model's baseline AI investment relative to the socially efficient level; a calibrated scenario, not an observed fact or forecast.

  8. 3,984

    historical individual-bank runs extracted and validated from US newspapers, 1863–1934.

  9. +80% to −20%

    estimated host-county productivity advantage in 1982 and disadvantage in 2010 after China's 1950s industrial-cluster program.

  10. −4.6 versus +2.7 percentage points

    the hospital experiment's intent-to-treat effects on any prescription and diagnostic testing; direction is observed, patient welfare is not.

07

Where economists disagree

A disagreement is useful when evidence can resolve it

01

Is the recent US productivity acceleration an AI efficiency gain?

Position A

Sectoral and cross-country correlations are early evidence of realized AI productivity. AI-exposed US sectors grew 3.7% versus 1.7% elsewhere; higher national AI use strengthens the sectoral relationship.

Position B

Since early 2024, higher utilization accounts for essentially all measured TFP growth; underlying efficiency shows little acceleration.

The balance favors “not yet proven.” Position A has useful heterogeneity but no causal adoption shock. Position B has a coherent accounting adjustment but relies on a coarse inferred utilization series subject to revision.

Resolution: Firm-level adoption timing linked to revenue, prices, quantities, hours, capital services, and intangible investment; persistent within-industry differences after utilization adjustment.

02

Is demographic decline a growth drag or an innovation stimulus?

Position A

Fewer young workers reduce labor supply, ideas, demand, and fiscal capacity; aging lowers innovation.

Position B

Labor scarcity raises wages and redirects invention and adoption toward labor-saving technology, offsetting the quantity loss.

The new paper materially shifts the balance against mechanical pessimism but not against fiscal concern. Per-worker productivity and pension arithmetic are different questions.

Resolution: Prospective quasi-experiments in regions facing predictable cohort contraction, with direct measures of automation investment, TFP, migration, prices, and fiscal transfers.

03

Are long-run inflation expectations anchored by credibility?

Position A

Low sensitivity to surprises—especially after formal targeting—reveals a credible nominal anchor.

Position B

The same aggregate pattern follows from lifetime learning about low inflation persistence; anchoring can unravel after repeated shocks.

Aggregate stability alone no longer discriminates. Age heterogeneity gives experience-based learning a distinctive empirical advantage, while not proving that credibility is irrelevant.

Resolution: Repeated shocks combined with panel expectations, information treatments, and cross-country target-regime variation that separates policy beliefs from experienced persistence.

04

Do high-tech clusters causally raise top-inventor productivity?

Position A

Larger clusters create knowledge spillovers; the original baseline elasticity was 0.0676.

Position B

Corrected mover and IV designs are statistically insignificant; even exogenously seeded clusters can reverse from large early productivity gains to long-run lock-in.

The original causal magnitude should not be used for policy calibration. A true short-run agglomeration benefit remains possible—Heblich et al. observe one—but persistence depends on entry, innovation, competition, and adaptability.

Resolution: Clean administrative replications with versioned code and exogenous cluster shocks, followed long enough to measure entry, innovation, markups, worker mobility, and productivity after the original anchor firms mature.

05

Should AI systems be optimized for benchmark accuracy?

Position A

More capable, more accurate models should weakly improve downstream decisions; guardrails protect users against high-cost errors.

Position B

Optimal informativeness depends on priors, verification, task stakes, and whose loss function the guardrails encode. More inference or greater caution can redirect rather than improve decisions.

Position B is currently more useful economically. Benchmark accuracy remains necessary evidence about capability, but it is not a sufficient objective for deployment. The hospital study identifies directional pass-through, not its welfare sign.

Resolution: Randomized deployment tests varying model behavior, verification cost, user expertise, and downstream loss, with final health, output, or welfare outcomes—not only answer accuracy or utilization.

08

Emerging research frontier

Seven questions the next editions must track

  1. 01

    The production economics of verification

    Which tasks have low verification cost, who should verify, and how does verification scale with output volume? This is likely to become the missing factor input in task-based AI models.

  2. 02

    Workflow-level randomized trials

    Most experiments randomize access for individuals. The frontier is organization-level treatment: process redesign, roles, incentives, data, training, and model use together, with value added and quality measured over at least a year.

  3. 03

    AI and the formation of junior human capital

    If AI removes entry-level drafting, coding, research, or diagnosis, does it accelerate learning through feedback or erode the practice that builds judgment? Education RCTs and workplace career panels should converge on this question.

  4. 04

    Utilization-adjusted AI productivity accounts

    Economists need timely accounts that separate infrastructure demand, capital deepening, utilization, intangible reorganization, quality change, and true technical efficiency at firm, sector, and aggregate levels.

  5. 05

    Endogenous AI exposure and rent incidence

    Task exposure must be joined to product-market competition, upstream concentration, bargaining, ownership of data, and the elasticity of model supply. This will determine whether gains reach workers, firms, providers, or consumers.

  6. 06

    Financial networks in the AI buildout

    Specialized collateral, long-term power contracts, leases, debt, cloud commitments, and circular equity ties could propagate a sectoral disappointment. The BIS calibration is a prompt for empirical balance-sheet mapping.

  7. 07

    Small AI and state capacity

    Courts provide one credible result. The next wave should test tax administration, health triage, agricultural extension, and benefits delivery in low-capacity states—where scarce expertise makes the social return potentially large but local-language data and institutional trust are binding.

09 · The Schymura Take

The shadow price of the bottleneck

The obvious AI investment case starts with expensive labor: automate the highest wage and capture the largest saving. July's evidence points somewhere less obvious.

The first high-return deployments may be where expert output is rationed rather than merely costly—courts with case backlogs, clinics with too few specialists, tax administrations with unprocessed files, schools without enough individualized feedback. In those systems, AI does not need to replace the expert to create value. It needs to expand the number of cases on which scarce judgment can be exercised, while routing the machine toward tasks that are cheap to verify.

That is an inference from the literature, not a finding any single paper establishes. But if it is right, the relevant investment metric is not wage exposure. It is the shadow price of the bottleneck multiplied by the verifiability of the delegated task.

This changes both business strategy and policy. A sophisticated chatbot added to an unconstrained workflow may save minutes that the organization cannot monetize. A less glamorous, locally adapted system added to a queue-constrained institution may create real capacity. The paradox is that AI's first large social productivity dividend may appear not where measured labor productivity is already high, but where institutional scarcity has kept valuable output from being produced at all.

10

Serious contributions screened

The complete 70-record research ledger

Scores guide editorial attention; they are not cardinal estimates of scientific quality.

Open source appendix
ScorePaperStatusTopicUseSource
92Who Pays for Unions?Dodini; Stansbury; WillénAccepted articleQJELabor/IOYes
91The Effect of Low-Skill Immigration Restrictions on US Firms and WorkersClemens; LewisPeer-reviewed articleAEJ: AppliedImmigration/laborNo
90Global Working HoursGethin; SaezAccepted articleQJELabor/measurementNo
87Subjective Earnings RiskCaplin; Gregory; Lee; Leth-Petersen; SæverudOnline-first articleReview of Economic StudiesLabor/householdsNo
93Is Growth Additive?Jones; López-Salido; PhilipponWorking paperNBER 35415GrowthYes
92Baby Busts and Growth BoomsAcemoglu; Autor; Beirne; ScottWorking paperNBER 35401Demography/growthYes
88Literature Review and Evidence AggregationGanong; Garg; KasyWorking paperNBER 35403MethodsRadar table
84How Does Monetary and Fiscal Policy Affect the Economy in the Face of Large Shocks?Kaplan; MiyaharaWorking paperNBER 35400MacroNo
83Importing Aggregate DemandLian; Mukhin; WolfWorking paperNBER 35402International macroNo
89The Use and Misuse of Average and Marginal Energy PricesHawkins-Pierot; WagnerWorking paperNBER 35406Energy/climateNo
90AI PremiumBorri; Liu; TsyvinskiWorking paperNBER 35451AI/financeYes
86Beyond Mean ExposureGhaddar; Babiarz; Li; Meng; Miller et al.Working paperNBER 35428Environment/healthNo
87Estimating the Economic Effects of Federally Funded R&DCampbell; Nelson; Schrag; Williams; WroblewskiInstitutional modelCBO WP 2026-08 / NBER 35472R&D/fiscalNo
86Training AI for When Humans Will Use ItBryan; GansTheory WPNBER 35490AI/informationFocus
92Directional AI Advice: Experimental Evidence from HealthcareChen; Li; Meng; Qiu; YangPreregistered field RCT WPSIEPR/arXivAI/healthYes
88A Model of the Data EconomyFarboodi; VeldkampOnline-first theory articleReview of Economic StudiesData/AINo
89Structural Change in Production Networks and Economic GrowthGaggl; Gorry; vom LehnGrowth-accounting/model articleReview of Economic StudiesNetworks/growthNo
85The Selective Disclosure of Evidence: An ExperimentFarina; Fréchette; Ispano; Lizzeri; PeregoLab experiment articleReview of Economic StudiesInformation/behaviorNo
90Seemingly Anchored Inflation ExpectationsMalmendier; NagelWorking paperNBER 35395 / BFIInflationDisagreement
86From Stocks to Flows: Debt Service and Fiscal SustainabilityEichengreen et al.Working paperNBER 35459Fiscal historyNo
85Financial Sanctions and the Global Payments NetworkMatvos; NeimanWorking paperNBER 35453Geoeconomics/financeNo
83Monetary-Policy Pass-Through in 96 CountriesAnagol; WangWorking paperNBER 35439MonetaryNo
88The AI Investment RaceRungcharoenkitkulInstitutional researchBIS WP 1367AI/financeFocus
86How Might Fiscal Policy Respond to the Rise of AI?Dynan; Elmendorf; SheinerScenario/policy WPNBER 35437AI/fiscalFocus
85Smooth Diagnostic ExpectationsBianchi; Ilut; SaijoOnline-first theory articleReview of Economic StudiesExpectations/macroNo
83Does Multi-Agent Debate Improve AI Feedback on Research Papers?Havránek; IrsovaPreregistered experimentCEPR 21752AI/methodsFocus
88Higher Utilisation Explains the Recent Surge in Productivity GrowthBoyle; Fernald; LiInstitutional analysisCEPR/VoxEUAI/productivityFocus
87The AI Buildout and the EconomySoto; Thieu; AllenInstitutional measurementFederal Reserve FEDS NoteAI/macroFocus
91Bank Runs With and Without Bank FailureCorreia; Luck; VernerSubstantial July revisionNBER 35504 / NY Fed 1198Banking/historyYes
93Industrial Clusters in the Long RunHeblich; Seror; Xu; ZylberbergOnline-first articleReview of Economic StudiesIndustrial policy/growthYes
92An Evaluation of Protected Area Policies in the European UnionGrupp; Mishra; Reynaert; van BenthemOnline-first articleReview of Economic StudiesEnvironment/policyTheme
90Organizational Incentives and the Returns to Technology AdoptionAdhvaryu; Gade; Gandhi; Molina; NyshadhamField RCT WPNBER 35445Organization/technologyTheme
84Supply Chain RiskCastro-Vincenzi et al.Working paperNBER 35496Trade/firmsNo
87Effects of Lottery Incentives for Influenza VaccinationMoran et al.Large RCT WPNBER 35537Health/behaviorTheme
92Courts of TomorrowMehmood; Goessmann; AshFormal circulation/substantial revision of field RCTCEPR 21783AI/public sectorYes
85Long-Run Effects of Universal Pre-Primary Education ExpansionBerlinski; Cruces; Galiani; Gertler; GonzalezWorking paperCEPR 21782EducationNo
87The Transmission of Reliable and Unreliable InformationGraeber; Noy; RothAccepted articleQJEBehavioral/informationNo
94The Effects of U.S. Public R&D on Global GrowthDe Souza; Fieldhouse; Mertens; Nath; RameyWorking paperNBER 35517Innovation/growthYes
88Who Is Afraid of Eurobonds?Bianchi; Fang; Melosi; Rogantini PiccoSubstantially revised quantitative-theory WPNBER 35510Euro/macroNo
85Information and Macroeconomic Expectations: Global EvidenceD'Acunto; WeberNBER-series issue of earlier BFI WPNBER 35511ExpectationsNo
85Immigration and Macroeconomic Outcomes in OECD CountriesBasso; Mathur; PeriWorking paperNBER 35523Migration/macroNo
89Capital Services in Global Value ChainsDingAccepted articleQJETrade/capitalNo
90The Class Gap in Career ProgressionStansbury; RodriguezPeer-reviewed articleEconometrica 94(4)Inequality/laborNo
89Mechanism Design for Personalized PolicyDizon-Ross; ZuckerPeer-reviewed articleEconometrica 94(4)Mechanism design/healthNo
88Experimental Evidence on the Learning Impact of Generative AIContractor; ReyesLab/field RCT WPIZA 18792AI/educationYes
80When AI Does the WorkNikolova; Milanova; WangSurvey experiment WPIZA 18784AI/job qualityFocus
79Let Me Check on YouNikolovaVignette RCT WPIZA 18782AI/job qualityFocus
81AI and Technological Relatedness in Regional InnovationD'Alessandro; Santarelli; VivarelliObservational WPIZA 18817AI/innovationNo
80Let's Chat: Leveraging Chatbot Outreach for Improved Course PerformanceMeyer et al.Preregistered WPNBER 35397Digital educationNo
87Do Markets Believe in Transformative AI?Andrews; FarboodiRevised working paperNBER 34243AI/asset pricingFocus
82Optimal Medical Liability for AIChanRevised theory WPNBER 35321 / HBSAI/health lawNo
89Ideological Alignment and Evidence-Based Policy AdoptionGarcía-Hombrados et al.Peer-reviewed articleAER 116(7)Political economyNo
92High-Tech Clusters and Top Inventors: CommentWiebePeer-reviewed commentAER 116(7)Replication/innovationYes
91The Aggregate Costs of Uninsurable Business RiskBoar; Gorea; MidriganPeer-reviewed article / structural modelAER 116(7)Entrepreneurship/macroNo
90Germs in the FamilyDaysal; Ding; Rossin-Slater; SchwandtPeer-reviewed articleAER 116(7)Health/human capitalNo
90Turbocharging Profits?Gupta; La Forgia; SacarnyPeer-reviewed articleAEJ: Applied 18(3)Health/IONo
89A Welfare Analysis of Policies Impacting Climate ChangeHahn; Hendren; Metcalfe; Sprung-KeyserPeer-reviewed articleAER 116(7)Climate/publicNo
91Personal Holding Companies, Tax Progressivity, and InequalityRing; Seim; ZucmanWorking paperNBER 35534Tax/inequalityYes
89Degrees of MobilityBurland et al.Field RCT WPNBER 35434Education/mobilityNo
86Physician Learning and ForgettingFitzgerald; Gupta; SchwabWorking paperNBER 35502Health/productivityNo
86Quality, Selection, and CongestionGowrisankaran; Marquardt; TownWorking paperNBER 35426Health/IONo
91Double Robustness of Local Projections and Some Unpleasant VARithmeticMontiel Olea; Plagborg-Møller; Qian; WolfPeer-reviewed articleEconometrica 94(4)Econometrics/macroNo
87Economic Growth and the Rise of Large FirmsChenPeer-reviewed articleEconometrica 94(4)Firms/growthNo
87Climate Policy in the Wide WorldHassler; Krusell; OlovssonPeer-reviewed lecture/articleEconometrica 94(4)Climate/macroNo
87Firm-to-Firm Trade: Imports, Exports, and the Labor MarketEaton; Kortum; KramarzPeer-reviewed articleEconometrica 94(4)Trade/firmsNo
85Macro Shocks and Firm Dynamics with Oligopolistic Financial IntermediariesVillaOnline-first structural articleReview of Economic StudiesFinance/macroNo
82AI and the US Economy: An Accounting PerspectiveCarpinelli; Natoli; TabogaAccounting researchInstitutional/SSRNAI/macroNo
88World Development Report 2026: The Promise of AIWorld Bank teamFlagship institutional reportWorld BankAI/developmentFocus
84AI Financial AdviceChoukhmane; de Silva; Lin; AkuzawaNew-series/revised WPNBER 35574AI/household financeNo
86AI Agents and Prompt Engineering in Econometric CodingGaliani; López; SosaBenchmark WPNBER 35588AI/researchFocus addendum

Verification record

Every paper in the main “Papers That Matter” section was checked against an original journal page, DOI, working-paper-series page, institutional repository, or author manuscript. Titles, author lists, series numbers, publication status, and dates were cross-checked at the source level. Numerical claims in the main narrative were retained only when located in the paper, abstract, table, figure, or official institutional summary.

The following interpretation rules were applied:

  1. An NBER, CEPR, IZA, university, central-bank, or BIS paper is described as a working paper or institutional research—not as peer reviewed.
  2. Model scenarios such as the BIS bust calibration and the Dynan–Elmendorf–Sheiner debt path are labeled scenarios, not forecasts or causal estimates.
  3. Correlations such as AI exposure versus sector productivity are not presented as causal.
  4. July series issuance is distinguished from an earlier manuscript date and from a later revision date.
  5. The World Bank report and August NBER items appear only in the focal-topic addendum permitted through the 12 August cutoff.
  6. No unresolved verification flag remains in the publishable narrative. Appendix entries with only qualitative summaries intentionally omit coefficients rather than rely on unverified secondary numbers.

Signal-score rubric

Scores use the requested weights: scientific rigor 25%, originality 20%, economic significance 20%, potential long-term importance 15%, policy/business relevance 10%, and surprise 10%. Differences of a few points should not be interpreted as statistically meaningful rankings.