Public R&D
Foreign TFP rises about 1% after a 1%-of-capital US R&D shock
National appraisal misses large global spillovers
July 2026 · Economics of AI
AI productivity is a conversion problem, not a capability statistic.
July moved the economic frontier from model capability to conversion: organizational design, verification, market structure, and institutional scarcity determine whether new technology becomes measured output.
Executive signal
July's research did not produce one dominant answer. It produced three unusually useful corrections.
The first concerns growth. Two ambitious papers challenge comfortable extrapolations from the past, but in opposite directions. Jones, López-Salido, and Philippon argue that US total factor productivity has historically advanced by roughly constant additions to its level, not by a constant exponential rate. That turns the familiar secular slowdown into something closer to an arithmetic implication. De Souza and coauthors, by contrast, estimate that US public R&D generates very large and slow-moving productivity spillovers abroad. Read together, the papers suggest that trend growth may be less automatic than standard models assume and more dependent on deliberately produced, internationally diffusing knowledge than national accounts reveal.
The second correction concerns scarcity. Acemoglu, Autor, Beirne, and Scott find that lower birth rates have been followed by faster growth in output per working-age adult, without a detectable loss of aggregate output. Their proposed mechanism is not a free demographic lunch. Scarcity of younger workers changes the direction of innovation and induces labor-saving investment. This is a more dynamic account of aging than merely subtracting workers from a fixed production function.
The third—and the month's clearest AI signal—concerns the unit of analysis. AI capability, tool adoption, task speed, worker output, firm value added, and aggregate productivity are not interchangeable variables. Pakistan's randomized JudgeGPT rollout raises completed cases by 6.3% only when access is paired with task-specific training. A hospital experiment shows that developer guardrails can redirect real prescriptions and tests without establishing whether health improves. A student experiment finds a 0.27-standard-deviation knowledge gain, but durable benefits are concentrated among users who ask AI to explain rather than to do the work. A Federal Reserve measurement note still sees a buildout phase, not broad transformation. And a competing aggregate interpretation attributes essentially all US TFP growth since early 2024 to higher utilization rather than better underlying efficiency.
The intellectual change is therefore subtle but consequential: the frontier has moved from asking what models can do to asking which production systems can convert those capabilities into verified economic output. The scarce input is increasingly not intelligence in the abstract. It is organizational design.
Across the broader radar, the same logic reappears. A policy label or a newly installed tool is not itself a treatment. Union incidence changes with market power; protected status without strict rules produces little visible environmental change; communication technology requires managerial incentives; and an industrial big push can become long-run lock-in. July was a month in which implementation stopped being a footnote and became the mechanism.
Economics Radar
Scores are editorial decision aids. Evidence labels remain more important than small rank differences.
Public R&D
National appraisal misses large global spillovers
Growth
Constant percentage-growth assumptions need justification
Industrial clusters
Anchor plants can crowd out renewal rather than seed it
Demography
Scarcity can redirect innovation
AI implementation
Organizational capital is part of the technology
AI guardrails
Developer objectives can propagate into real expert decisions
Replication
Place-based policy loses a key calibration
Inequality
Legal realization distorts progressivity and inequality
Inflation
Stability may unravel after repeated shocks
AI markets
Incidence is endogenous to use and market structure
AI learning
Human-capital effects depend on use mode
AI investment
Private races can create financial fragility
Evidence synthesis
Literature means can greatly overstate treatment effects
Frontier observatory
The charts describe this edition, not an imagined economy: signal scores, evidence labels, paper rankings, and the screened research ledger.
Public R&D
The papers that matter
Each record separates status, method, result, relevance, and limitation.
Gustavo De Souza (Federal Reserve Bank of Chicago), Andrew J. Fieldhouse (Texas A&M Mays Business School), Karel Mertens (Federal Reserve Bank of Dallas), Ishan B. Nath (Harvard Kennedy School and NBER), and Valerie A. Ramey (Stanford, Hoover Institution, and NBER).
A shock equal to 1% of the US federal R&D capital stock raises foreign TFP by approximately 1% after 12 years. The peak response is about 1.8% in non-OECD economies and 0.7% in OECD economies.
Callum J. Jones and David López-Salido (Federal Reserve Board); Thomas Philippon (New York University and NBER).
Additive TFP describes approximately 90 years of US evidence better than the geometric benchmark. Postwar TFP adds roughly 2.5% of its 1947 level annually, implying a declining percentage growth rate as the level rises.
Daron Acemoglu, David Autor, and Keelan Beirne (MIT); Andrew Scott (London Business School and Ellison Institute of Technology).
Lower birth rates predict higher GDP per working-age adult and wages, with no detectable fall in aggregate GDP or earnings. The evidence points to labor-saving technological change.
Samuel Dodini (Federal Reserve Bank of Dallas), Anna Stansbury (MIT Sloan), Alexander Willén (Norwegian School of Economics).
In the average private firm, higher unionization raises labor cost but reduces employment, output, and profit; labor's share does not rise. In manufacturing and less competitive markets, wages, employment, and output rise while markdowns fall.
Michael Wiebe (independent economist).
The original 0.0676 baseline elasticity is not supported by corrected mover-event-study and IV estimates, which are statistically insignificant. Errors include an omitted, correctly interacted event-time-zero treatment and an IV differenced across cities after unsorted data.
Sergio A. Correia (Federal Reserve Bank of Richmond), Stephan Luck (Federal Reserve Bank of New York), and Emil Verner (MIT Sloan and NBER).
Runs hit strong and weak banks, but weak pre-run fundamentals predict failure and coincide with larger local lending and real-activity contractions. Bank fundamentals alone predict runs with an AUC of about 0.69; adding aggregate and local information raises it to 0.80.
Marius A. K. Ring (University of Texas at Austin), David Seim (Stockholm University), and Gabriel Zucman (Paris School of Economics and UC Berkeley).
Around half of top-0.1% income is retained in PHCs; the highest-income owners distribute only 15–20% over 20 years. Effective total tax rates fall from around 50% in the upper middle to 15–20% at the top of wealth.
Sultan Mehmood (New Economic School); Christoph Goessmann and Elliott Ash (ETH Zurich; current manuscript affiliations).
At median district exposure, AI plus targeted training produces 1,848 additional resolved cases per district-year, 6.3% above the mean. Appeals slightly decline; measured writing quality does not deteriorate. Targeted instruction increases use at least fourfold relative to generic training and directs it toward bounded, verifiable tasks.
Nicola Borri (LUISS), Yukun Liu (University of Rochester, Simon Business School), and Aleh Tsyvinski (Yale and Cowles Foundation).
A value-weighted high-minus-low AI-beta portfolio earns 64.1 basis points per week (t=2.84), about 56 after standard factor controls (t=2.41). Interaction and communication skills load positively; analytical and information-use skills load negatively. Established engineering exposure measures explain under 2% of cross-firm beta variation.
Zara Contractor and Germán Reyes, Middlebury College; Reyes also IZA@LISER.
AI access raises immediate knowledge scores 6.7 percentage points from a 56.3% control mean, equal to 0.27 standard deviations. One week later, the gain is 5.1 points—76% of the immediate effect. Augmentation users retain gains; automation users' assisted essay gains disappear.
Stephan Heblich (University of Toronto and NBER), Marlon Seror (UQAM), Hao Xu (China Construction Bank), and Yanos Zylberberg (University of Bristol and CEPR).
Host counties were about 80% more productive than controls in 1982 but 20% less productive by 2010. Early industrialization raised non-agricultural household registration by roughly 30 percentage points, yet later entrants were low-productivity, innovation was limited, and markups were high.
Yuyu Chen (Peking University), Hongbin Li and Lingsheng Meng (Stanford University), and Xinyao Qiu and Qingxu Yang (University of Hong Kong).
Randomized access lowers the probability of any prescription by 4.6 percentage points from an 87% baseline and raises diagnostic testing by 2.7 points from a 23% baseline. Total spending is unchanged. The chatbot cautions against medication but usually recommends diagnostic testing, and clinical choices move in the same direction.
Focus · Economics of AI
Capability, adoption, task performance, organizational output, and aggregate productivity are different variables.
Strong evidence of technical progress; weak mapping to whole jobs
Strong on direction, moderate on comparable levels
Strong in bounded tasks; heterogeneous beyond them
Moderate-to-strong in selected occupations
Strong on decision direction; welfare remains unidentified
Moderate; causal value-added evidence remains scarce
Early and disputed
Early; realized wage and incidence effects remain thin
By June 2026, five propositions had become reasonably defensible.
First, frontier capability was improving rapidly, especially on bounded language, coding, and analytical tasks. Second, access to generative AI often raised task-level speed or quality in experiments. Third, adoption was broad but shallow: many workers had tried the tools, while only a small share of total hours and complete workflows used them. Fourth, effects were heterogeneous across the “jagged frontier”: AI could help weaker performers on well-defined tasks and harm performance where contextual judgment or verification exceeded user skill. Fifth, aggregate productivity and employment data did not yet contain an unambiguous AI break.
This left a translation puzzle. The literature could show impressive capability and credible task gains but offered much thinner evidence on worker output over time, firm value added, organizational redesign, and economy-wide TFP. July's contribution is not to close that gap. It is to identify the conversion mechanisms inside it.
The JudgeGPT experiment does something most AI evaluations do not: it randomizes not only access but the complement that tells users where and how to deploy it. Targeted training increases usage at least fourfold relative to generic technology training, moves prompts away from costly-to-verify open legal questions, and produces a 6.3% output increase without a measured quality loss. Access alone is not the relevant treatment.
This finding aligns with the non-AI garment-factory RCT in the broad radar. A communication technology creates no value until HR incentives change. The common mechanism is organizational: the effective productivity of a tool depends on whether someone owns adoption, workers know the tool's comparative advantage, and verification is assigned to tasks where it is economical.
Chen et al. randomize patient access to a medical chatbot before outpatient visits. Only 17.1% of those offered access use it, yet the intent-to-treat effects move real care: prescriptions fall 4.6 percentage points from an 87% baseline, while diagnostic testing rises 2.7 points from a 23% baseline. Total spending is unchanged.
The conversation logs reveal the mechanism. The chatbot attaches cautions to 90.8% of traditional-Chinese-medicine mentions and 87.6% of antibiotic mentions, but gives a clean recommendation in 94.5% of diagnostic-test discussions. The system's defensive, liability-sensitive design propagates into clinical choices, especially when physicians are receptive to patient input. This is not evidence that AI improved health: the paper has no clinical outcome with which to value fewer drugs against more tests. It is strong evidence that the developer's implicit loss function can become an input into another institution's production process.
Contractor and Reyes distinguish assisted output from later unaided performance. Their 0.27-standard-deviation immediate knowledge gain retains 76% after one week. But use logs reveal the boundary: students who ask for explanations preserve benefits; students who outsource text see short-run essay gains vanish when the tool is removed. AI can be a tutor or a substitute for practice. “Used AI” is too coarse a variable to tell which.
AI Premium uses 380 trillion tokens of actual consumption to infer an AI factor and firm-level market betas. Technical exposure measures explain under 2% of their cross-firm variation. Market exposure is more positive for interaction and communication content and more negative for analytical and information-use skills. That is surprising only if capability maps are mistaken for equilibrium incidence. Prices reflect adoption, complementarity, competition, bargaining, and expected rents—not merely whether an LLM can perform a task.
Havránek and Irsova's negative experiment matters beyond academic feedback. Authors prefer a single frontier-model pass to two multi-agent debate systems, although one uses around 30 times the tokens. AI judges almost always rank the real human referee report last. The result is narrow but sharp: inference-time complexity can lower user value, and model-based evaluation can reverse the human outcome criterion.
The BIS AI Investment Race calibrates a dynamic winner-take-most contest among hyperscaler–lab coalitions. Its baseline produces investment at 1.51 times the efficient level, a 50% bust probability, and expected destruction of $189 billion per year. These are model outputs, not forecasts. Still, the mechanism is important: private investment can be individually rational and collectively excessive when firms race for a few dominant positions, finance specialized assets with debt, and own circular stakes in one another.
The macro evidence is not yet internally consistent—and should not be made to look so.
Scott Davis reports that the three most AI-exposed US sectors grew productivity by 3.7% annualized since early 2024 versus 1.7% elsewhere. They account for 16% of hours but 40% of US productivity gains. Across European countries, greater Claude use is associated with a stronger exposure–productivity slope. This is suggestive of realized AI gains, but explicitly correlational.
Boyle, Fernald, and Li offer a competing decomposition. US output per hour grew about 2.5% annually from 2023 to 2026Q1, one point above 2005–19. Measured TFP contributed 0.8 point to the acceleration, yet since early 2024 their utilization estimate accounts for essentially all TFP growth, leaving little utilization-adjusted improvement. In this reading, AI may have raised demand and uncertainty, causing firms to work existing inputs harder—not yet smarter.
Both can be true. AI-exposed sectors may be improving while aggregate TFP is dominated by cyclical intensity; or exposure may proxy for pre-existing sectoral advantages. The decisive evidence will be sustained, utilization-adjusted productivity divergence within industries after comparable adoption, ideally linked to firm-level workflows and value added.
Benchmarks assign equal or arbitrary weight to tasks. GDP weights outputs by market value; welfare also includes time, quality, variety, consumer surplus, and nonmarket services. Saturating an exam does not reveal which production bottleneck has been relaxed.
“Uses AI” can mean one query per month or a redesigned production process. The Fed note finds uptake rising with firm size but emphasizes shallow intensity. Surveys need task shares, workflow penetration, model costs, verification time, and complementary investment.
Courts can count cases, but justice quality is multidimensional. Education can count test answers, but durable reasoning differs from polished text. In many services, nominal revenue is deflated with imperfect prices, so free or cheaper AI-enhanced quality can disappear from measured real output.
Data centers and computers enter investment; imported equipment subtracts through net exports; software is partly capitalized; data, model fine-tuning, training, and workflow redesign are incompletely measured intangibles. The Fed estimates that AI-related components contributed meaningfully to GDP growth from 2025 to 2026Q1, while imports offset much of gross investment in some quarters.
Engineering scores ask what AI could perform. Usage logs ask what people currently delegate. Stock betas ask which cash flows or discount rates covary with AI demand. These measures should disagree. Treating one as ground truth creates false contradictions.
July's bank-run archive shows the upside of LLM extraction. But AI Premium, AI-judged paper feedback, text-based court-quality scores, and generated covariates all place model error inside the estimator. Validation must be designed around the target coefficient, not generic classification accuracy.
Fewer prescriptions, more diagnostic tests, more resolved cases, or more books are intermediate outcomes. Their value depends on clinical health, legal accuracy, consumer use, and the counterfactual cost of resources. AI research will overstate progress if it counts directional behavior as improved welfare without measuring the endpoint that gives the behavior economic value.
Numbers worth remembering
additional court cases resolved at median district exposure to JudgeGPT with targeted training, not AI access in isolation.
immediate knowledge gain from AI access in the Middlebury experiment; 76% persists one week later.
estimated foreign TFP response to a US R&D appropriations shock equal to 1% of federal R&D capital.
minimum reported posterior probability on the additive TFP model by 2022 across prior choices.
higher GDP per working-age adult associated with a one-point lower birth rate across countries over 1970–2020.
value-weighted high-minus-low AI-beta return in AI Premium; the short sample makes this a signal, not a settled premium.
BIS model's baseline AI investment relative to the socially efficient level; a calibrated scenario, not an observed fact or forecast.
historical individual-bank runs extracted and validated from US newspapers, 1863–1934.
estimated host-county productivity advantage in 1982 and disadvantage in 2010 after China's 1950s industrial-cluster program.
the hospital experiment's intent-to-treat effects on any prescription and diagnostic testing; direction is observed, patient welfare is not.
Where economists disagree
Sectoral and cross-country correlations are early evidence of realized AI productivity. AI-exposed US sectors grew 3.7% versus 1.7% elsewhere; higher national AI use strengthens the sectoral relationship.
Since early 2024, higher utilization accounts for essentially all measured TFP growth; underlying efficiency shows little acceleration.
The balance favors “not yet proven.” Position A has useful heterogeneity but no causal adoption shock. Position B has a coherent accounting adjustment but relies on a coarse inferred utilization series subject to revision.
Resolution: Firm-level adoption timing linked to revenue, prices, quantities, hours, capital services, and intangible investment; persistent within-industry differences after utilization adjustment.
Fewer young workers reduce labor supply, ideas, demand, and fiscal capacity; aging lowers innovation.
Labor scarcity raises wages and redirects invention and adoption toward labor-saving technology, offsetting the quantity loss.
The new paper materially shifts the balance against mechanical pessimism but not against fiscal concern. Per-worker productivity and pension arithmetic are different questions.
Resolution: Prospective quasi-experiments in regions facing predictable cohort contraction, with direct measures of automation investment, TFP, migration, prices, and fiscal transfers.
Low sensitivity to surprises—especially after formal targeting—reveals a credible nominal anchor.
The same aggregate pattern follows from lifetime learning about low inflation persistence; anchoring can unravel after repeated shocks.
Aggregate stability alone no longer discriminates. Age heterogeneity gives experience-based learning a distinctive empirical advantage, while not proving that credibility is irrelevant.
Resolution: Repeated shocks combined with panel expectations, information treatments, and cross-country target-regime variation that separates policy beliefs from experienced persistence.
Larger clusters create knowledge spillovers; the original baseline elasticity was 0.0676.
Corrected mover and IV designs are statistically insignificant; even exogenously seeded clusters can reverse from large early productivity gains to long-run lock-in.
The original causal magnitude should not be used for policy calibration. A true short-run agglomeration benefit remains possible—Heblich et al. observe one—but persistence depends on entry, innovation, competition, and adaptability.
Resolution: Clean administrative replications with versioned code and exogenous cluster shocks, followed long enough to measure entry, innovation, markups, worker mobility, and productivity after the original anchor firms mature.
More capable, more accurate models should weakly improve downstream decisions; guardrails protect users against high-cost errors.
Optimal informativeness depends on priors, verification, task stakes, and whose loss function the guardrails encode. More inference or greater caution can redirect rather than improve decisions.
Position B is currently more useful economically. Benchmark accuracy remains necessary evidence about capability, but it is not a sufficient objective for deployment. The hospital study identifies directional pass-through, not its welfare sign.
Resolution: Randomized deployment tests varying model behavior, verification cost, user expertise, and downstream loss, with final health, output, or welfare outcomes—not only answer accuracy or utilization.
Emerging research frontier
Which tasks have low verification cost, who should verify, and how does verification scale with output volume? This is likely to become the missing factor input in task-based AI models.
Most experiments randomize access for individuals. The frontier is organization-level treatment: process redesign, roles, incentives, data, training, and model use together, with value added and quality measured over at least a year.
If AI removes entry-level drafting, coding, research, or diagnosis, does it accelerate learning through feedback or erode the practice that builds judgment? Education RCTs and workplace career panels should converge on this question.
Economists need timely accounts that separate infrastructure demand, capital deepening, utilization, intangible reorganization, quality change, and true technical efficiency at firm, sector, and aggregate levels.
Task exposure must be joined to product-market competition, upstream concentration, bargaining, ownership of data, and the elasticity of model supply. This will determine whether gains reach workers, firms, providers, or consumers.
Specialized collateral, long-term power contracts, leases, debt, cloud commitments, and circular equity ties could propagate a sectoral disappointment. The BIS calibration is a prompt for empirical balance-sheet mapping.
Courts provide one credible result. The next wave should test tax administration, health triage, agricultural extension, and benefits delivery in low-capacity states—where scarce expertise makes the social return potentially large but local-language data and institutional trust are binding.
09 · The Schymura Take
The obvious AI investment case starts with expensive labor: automate the highest wage and capture the largest saving. July's evidence points somewhere less obvious.
The first high-return deployments may be where expert output is rationed rather than merely costly—courts with case backlogs, clinics with too few specialists, tax administrations with unprocessed files, schools without enough individualized feedback. In those systems, AI does not need to replace the expert to create value. It needs to expand the number of cases on which scarce judgment can be exercised, while routing the machine toward tasks that are cheap to verify.
That is an inference from the literature, not a finding any single paper establishes. But if it is right, the relevant investment metric is not wage exposure. It is the shadow price of the bottleneck multiplied by the verifiability of the delegated task.
This changes both business strategy and policy. A sophisticated chatbot added to an unconstrained workflow may save minutes that the organization cannot monetize. A less glamorous, locally adapted system added to a queue-constrained institution may create real capacity. The paradox is that AI's first large social productivity dividend may appear not where measured labor productivity is already high, but where institutional scarcity has kept valuable output from being produced at all.
Serious contributions screened
Scores guide editorial attention; they are not cardinal estimates of scientific quality.
Every paper in the main “Papers That Matter” section was checked against an original journal page, DOI, working-paper-series page, institutional repository, or author manuscript. Titles, author lists, series numbers, publication status, and dates were cross-checked at the source level. Numerical claims in the main narrative were retained only when located in the paper, abstract, table, figure, or official institutional summary.
The following interpretation rules were applied:
Scores use the requested weights: scientific rigor 25%, originality 20%, economic significance 20%, potential long-term importance 15%, policy/business relevance 10%, and surprise 10%. Differences of a few points should not be interpreted as statistically meaningful rankings.