Growth theory · machine economy · artificial intelligence

The Token Is Not the Factor

What Solow gets right about AI, and why the productivity question begins where the benchmark ends.

AI looks weightless at the interface and industrial underneath. The economic unit is not a token but a verified outcome, produced jointly by compute, energy, data, human judgment, and organizational capital.

AuthorDr. Michael Schymura
Published2 August 2026 · 42 min
Research cut-off2 August 2026

Robert Solow, 2008. Olaf Storbeck, CC BY-SA 2.0.

Executive synthesis

The factory behind the cursor

In 1956, Robert Solow published a model that separated the accumulation of capital from technical change.[1] The distinction looked austere even then: output, capital, effective labour, saving, depreciation. Seventy years later, when an answer can arrive before the user has lifted a hand from the keyboard, that austerity has become useful again. It asks us to classify before we celebrate, to measure before we extrapolate, and to distinguish a richer economy from an economy whose growth rate has permanently changed.

There is a data centre somewhere at the end of every elegant prompt. The sentence appears instantly, with no loading dock, factory whistle, or visible inventory — only a cursor, a reply, and the impression that intelligence has become a utility so light that the old economic categories no longer apply. Yet behind the interface sit structures, accelerators, memory, networks, electricity, cooling, software, data rights, human reviewers, and organizations that decide whether the answer may enter a production process at all. The interface is weightless; the production system is not.

Then follow the metaphors. Tokens are the new labour. Compute is the new oil. Models are the new capital. Data are the new land. Intelligence is a factor of production. Each phrase illuminates one surface and blurs the mechanism below it. None is precise enough to run a production account, allocate capital, or identify a causal productivity effect. A vivid analogy is not yet an economic object.

That proposition has eight immediate implications. Tokens are activity counts rather than homogeneous factors. Training and inference belong to different economic categories. Capital deepening can generate a very large level effect without creating permanent growth. Task acceleration can vanish inside system bottlenecks. The Solow residual becomes unusually treacherous when hidden organizational investment and utilization move together. AI capital depreciates in layers. Cheap inference can increase total spending and electricity demand. And the empirical frontier has shifted from capability to translation. Read as a programme, the question is no longer whether a model can perform a task. It is whether the task becomes an accepted, durable, valuable outcome.

Figure · economic classification

The machine is a production chain, not a new letter.

Structures

Data centres, cooling, grid connections

Capital stock

Compute

Accelerators, memory, networks, software

Capital service

Energy

Electricity and cooling consumed in use

Intermediate flow

Data + evaluations

Coverage, permissions, tests, feedback

Produced complement

Quality-adjusted machine service

Attempts × task value × success × acceptance × durability

Raw tokens sit inside attempts. Reliability, human verification, and downstream acceptance determine whether an attempt becomes a service.

Economic outcome

A release shipped. A case resolved. A decision improved.

The meaningful unit appears at the end of the production system, after review, integration, recovery, and demand.

Classification follows gross-output and capital-service logic. Ownership determines the boundary: inference is output for a provider, an intermediate purchase for a customer, and an internal capital service for an owner-operator.

Section 01

The seductive error of counting what is easy

Tokens are countable, frequent, and available in dashboards. That makes them irresistible. Gross domestic product arrives late. Productivity arrives later. Firm outcomes are proprietary. A token counter updates before the model finishes its sentence. But ease of measurement does not create an economic unit, and observability does not make a quantity homogeneous.

One token can be part of a correct answer, a hallucination, a duplicated context window, a cached result, a rejected draft, or a reasoning trace that prevents a million-euro error. Input and output tokens have different prices. Different models produce different quality from the same number. Some systems become more verbose as they become more capable. Agentic systems call tools and one another, multiplying tokens before producing a single accepted action. The count measures movement through a machine, not the value produced at the end.

Treating tokens as a primary factor also creates double counting. The token is produced by a model running on accelerators, memory, and networking, using electricity and systems software. Counting chips, energy, cloud expenditure, and tokens as four independent primary inputs makes one production chain appear several times. The accounting becomes a hall of mirrors precisely when it should become more disciplined.

The defensible unit is harder: a quality-adjusted, verified machine service. In operational terms, the account records whether the invoice was reconciled, the release shipped, the customer case resolved, the diagnosis improved, the experiment selected, or the decision changed. It then records whether the outcome survived review and what the full cost of producing it was. The answer belongs in the production account; the token count belongs in the telemetry.

The word “verified” carries economic weight. AI turns uncertainty into a production input. A model can produce a plausible answer cheaply and an accountable answer expensively. Human review, evaluation infrastructure, auditability, recovery design, and expected error loss are not governance outside the system. They are inputs inside it. Bryan and Gans formalize the implication: when humans decide whether to inspect, combine, or act on predictions, maximizing unconditional model accuracy need not maximize decision value.[2] Read as: a slightly less accurate model can create more economic value when its errors are legible and cheaply checked.

Outcome cost

Inference + human review + rework + expected error loss + integration overhead

That is the marginal cost a firm should compare with the marginal value of the outcome. The benchmark ends where the organization begins, and the economic measurement must continue through the organization or not claim productivity at all.

Section 02

Training is a stock; inference is a flow

Training creates or improves a model. Economically, the expenditure can produce an intangible stock: weights, methods, evaluations, data pipelines, and the organizational capacity to train again. The stock yields services over time and loses value when better models and hardware arrive, data drift, licences expire, or evaluation sets cease to describe the production environment. Training belongs in an accumulation equation.

Inference is different. It is the service flow generated when that stock is deployed with hardware, energy, memory, networks, and software. A cloud customer buys the flow. A provider combines inputs to sell it. An owner-operator consumes capital services internally. One activity builds productive capacity; the other uses capacity to produce an intermediate or final service. Collapsing both into “AI spend” obscures the very distinction that growth theory was designed to preserve.

The accounting boundary therefore decides what AI “is.” At the provider, inference is gross output produced with capital, labour, electricity, and purchased services. At the client, purchased inference is an intermediate input. If the client owns the accelerator, the service enters through capital. If employees build durable evaluations and workflows but their time is expensed, part of the new organizational asset disappears from measured investment. The technology has not changed; the boundary has.

Figure · conceptual model

AI capital depreciates on several clocks.

Accounting life is not economic life. The relevant measure is the decline in the present value of service, including obsolescence, migration, and revalidation costs.

Data-centre shell
DecadesPhysical wear dominates
Power + cooling
Years to decadesLocation and utilization matter
Accelerator system
YearsCost competitiveness can vanish first
Frontier model
Months to yearsEconomic obsolescence exceeds physical wear
Evaluation + workflow
Design-dependentPortable systems can outlive models
Relative bars are explanatory, not empirical duration estimates. Actual service lives depend on utilization, architecture, prices, model compatibility, portability, and the decision environment.

AI has at least four clocks. Buildings age slowly. Accelerators age faster. Frontier models age faster still. Workflows can age either quickly or slowly, depending on whether they are bound to one model or built around portable evaluations, permissions, and decision rules. Physical life and economic life diverge because a server may still run long after it has ceased to be cost-competitive.

This yields a strategic implication that appears mundane until model churn accelerates. The durable asset may not be the model. It may be the organization’s capacity to replace the model without losing control of the production process. Model-agnostic orchestration, evaluation suites, failure taxonomies, and explicit decision rights lower effective depreciation because they preserve service when the visible technology changes.

Caplin’s concept of “planning capital” extends the point.[22] Investment creates capacity, but operation also creates diagnostic knowledge about where capacity is useful, where it fails, and what the next investment should be. In enterprise AI, the map of eligible workflows, exceptions, costs, latency, migrations, and realized outcomes is itself a capital stock. Much of it is produced inside operating expense, which is exactly why conventional measurement misses it.

Section 03

An investment boom is not yet a growth regime

In its cleanest form, the Solow model says that aggregate output is produced by capital and effective labour. Constant returns allow us to divide by effective labour and study capital per effective worker. The law of motion is familiar; its warning is not.

Aggregate output from capital and effective labour
(1)
Yt=F ⁣(Kt,AtLt)Y_t = F\!\left(K_t,\, A_t L_t\right)

A scales labour effectiveness; it is not a miscellaneous bucket for every unmeasured AI input.

Law of motion for capital per effective worker
(2)
k˙t=sf(kt)(n+g+δ)kt\dot{k}_t = s f(k_t) - \left(n + g + \delta\right)k_t

Saving builds the stock; population growth, technical progress, and depreciation dilute it.

A higher saving or investment rate moves the economy toward a higher steady-state level of capital and output per effective worker. Growth is faster during the transition. Once the new position is reached, output per worker again grows at the rate of labour-augmenting technical progress. The economy can become substantially richer without permanently accelerating.

This is the cleanest way to interpret the data-centre build-out. It can be historic and still be a level transition in the Solow sense. More accelerators per worker can raise output. Better kernels and model architectures can increase the service from every accelerator. Faster research can improve the production of ideas. Autonomous research could eventually feed back into the rate of technical progress. These are distinct mechanisms, and a forecast that merges them into one “AI effect” loses the causal structure.

The official U.S. accounts show the investment wave before they show an unambiguous AI productivity wave. The April 2026 BEA–BLS integrated production account attributes, over 2021–2024, 0.29 percentage points of average annual aggregate value-added growth to software capital, 0.24 points to R&D capital, and 0.60 points to integrated TFP.[3]Read as: software, research, and the residual all matter, but the decomposition is not a causal estimate of generative AI.

BLS reports that private-business software investment grew at an annual rate of 11.1% in the 2019–2024 business cycle, faster than in the previous cycle.[4] Read as: capital deepening is visible and substantial. What remains unidentified is how much of the expenditure has changed production recipes rather than equipped old recipes with more expensive tools.

General-purpose technologies also require complementary investment in skills, standards, infrastructure, products, and organizational routines. Much of that effort is expensed, depressing measured productivity while a hidden stock accumulates. Later, when the stock produces, output rises without a corresponding measured input. The productivity J-curve describes this sequence.[5] AI may create an unusually steep curve because access arrives in an afternoon while permissions, evaluations, process redesign, training, exception handling, and trust arrive on institutional time.

Interactive model · simulated evidence

Build capital. Then confront the weak links.

The model separates AI accumulation from the organizational, verification, and energy constraints that determine whether machine potential becomes economic output.

Simulation scenario

Figure 1 · output per worker

Transition path, index at year 0 = 100

AI systemConventional path
Simulated output per worker over fifty yearsThe accent line shows the selected AI system scenario. The dashed neutral line shows a conventional-capital counterfactual. The horizontal axis runs from year zero to year fifty. The vertical output index runs from zero to 300. Exact values are available through the year control and table below.07515022530001020304050Years after investment shock
Pointer or arrow keys
Illustrative discrete-time simulation. Verification, organization, and energy attenuate machine services multiplicatively; values are model outputs, not forecasts.

Assumptions

8%

Output reinvested in compute, models, and AI equipment.

65%

Share of machine capability surviving review, integration, and demand.

22%

Economic obsolescence of hardware, models, and embedded workflows.

3.0%

Evaluations, permissions, process redesign, skills, and decision rights.

1.4

Normalized capacity for power, cooling, networks, and grid access.

2.5%

Annual quality gain in machine services at a fixed AI capital stock.

32.1%

50-year output lift

against conventional path

1.7%

Final annual growth

level and technical progress

20.1%

Machine potential realized

quality-adjusted service

Organization

Tightest complement

51.1%

Model assumptions and text fallback

Output combines conventional capital with quality-adjusted machine service. AI and organizational stocks accumulate through investment and depreciate independently. Population and labor-augmenting technology dilute capital per effective worker. The simulation deliberately omits prices, strategic competition, distribution, and general equilibrium, so it is a mechanism demonstrator rather than a calibrated forecast.

Selected simulation values by year
YearOutput indexAI capitalRealized serviceTranslation gap
0100.00.200.090.55
10129.40.510.270.79
20163.00.580.401.00
30205.30.630.571.27
40247.10.660.611.79
50292.80.680.622.47

Section 04

The residual contains our ignorance

Growth accounting often decomposes output growth into measured capital, measured labour, and a residual called total factor productivity. Under restrictive assumptions, the residual can be interpreted as technical efficiency. In applied work, it also contains what the statistician did not measure well: utilization, effort, output quality, new varieties, intangible investment, reallocation, markups, spillovers, and error.

Growth accounting decomposition
(3)
ΔlnY=αΔlnK+(1α)ΔlnL+ΔlnA\begin{aligned}\Delta\ln Y &= \alpha\,\Delta\ln K + (1-\alpha)\,\Delta\ln L \\ &\quad + \Delta\ln A\end{aligned}

The final term is calculated after measured input contributions are removed; interpretation requires stronger assumptions than arithmetic.

Consider a firm that buys an AI licence, assigns senior staff to create evaluation sets, slows production while it redesigns workflows, and records almost all of that effort as current expense. Measured output can fall while a hidden capital stock rises. Later, when the workflow stabilizes, output can increase without a measured input. The same project first depresses measured productivity and then flatters the residual. The residual did not discover the mechanism; it absorbed it.

The timing of the recent U.S. productivity acceleration invites an even simpler story: ChatGPT arrived, measured TFP rose, therefore AI caused the increase. Boyle, Fernald, and Li examine output per hour growth of roughly 2.5% annually from early 2023 through early 2026, about one percentage point above the 2005–2019 pace.[6] After adjusting for more intensive use of existing labour and capital, they attribute essentially all of the exceptional acceleration to utilization. Read as: measured TFP rose, but the coincidence does not identify technical change.

That result does not prove that AI has produced no efficiency gain. It proves that a residual cannot carry the causal burden by itself. AI investment may increase demand for construction, equipment, power, and services, while uncertainty causes other firms to delay irreversible hiring and capital decisions. Existing inputs are worked harder, output rises, and measured inputs lag. The residual receives the credit whether or not the production frontier moved.

Section 06

An economy of complements

AI commentary often asks whether the technology substitutes for labour or complements it. The answer can be yes to both, because the relation changes by layer. Within a task, a model may replace drafting time. Across tasks, the drafting system complements review and accountable judgment. Across the organization, it complements data, software, security, and management. In the labour market, it can substitute for one skill while increasing demand for another. One elasticity cannot describe a nested production system.

Final output with conventional capital, organization, and complementary tasks
(4)
Yt=AtKC,tαOtω×[01xj,tσ1σdj]σσ1\begin{aligned}Y_t &= A_t K_{C,t}^{\alpha} O_t^{\omega} \\ &\quad\times\left[\int_0^1 x_{j,t}^{\frac{\sigma-1}{\sigma}}\,\mathrm{d}j\right]^{\frac{\sigma}{\sigma-1}}\end{aligned}

When sigma is below one, tasks are complements: accelerating one stage raises the relative scarcity of the others.

Task production from human and machine services
(5)
xj,t=[βjHj,tηj1ηj+(1βj)Sj,tηj1ηj]ηjηj1\begin{aligned}x_{j,t} &= \Biggl[\beta_j H_{j,t}^{\frac{\eta_j-1}{\eta_j}} \\ &\quad +(1-\beta_j)S_{j,t}^{\frac{\eta_j-1}{\eta_j}}\Biggr]^{\frac{\eta_j}{\eta_j-1}}\end{aligned}

Each task has its own substitution elasticity eta; the production system need not share it.

The distinction matters. A high within-task elasticity, eta, says that machine service can replace human input in drafting a standard document. A low across-task elasticity, sigma, says that the document remains complementary to approval, consent, physical action, trust, or delivery. Labour substitution and system complementarity are not rival narratives. They are statements about different levels of aggregation.

The new task literature adds a third margin: simplification. Althoff and Reichardt model workers with multidimensional skills who choose occupations and accumulate capability on the job.[9] AI can lower the skill threshold required for a task, allowing a wider set of workers to perform it. In their scenarios, average wages rise and inequality narrows. Read as: AI need not only replace or augment; it can change who is eligible to produce.

The optimistic mechanism is conditional. If AI simplifies advanced tasks while preserving the learning path, it can broaden access to capability. If it removes the junior tasks through which workers acquire judgment, it can reduce the future supply of experts. The short-run output effect and the long-run human-capital effect can therefore have opposite signs. A production model that omits learning will declare victory before the relevant stock has depreciated.

Furthermore, verification is endogenous. The value of a prediction depends on what the human will do with it, which errors the human can recognize, and how the organization responds when both fail. Reliability is therefore a production parameter, not an ethics appendix. A rare error in a draft may have low economic weight; a rare error in a payment, medical recommendation, or access-control workflow can dominate expected value. Weak links make upside slow and downside fast.

Section 07

The physical economy behind intelligence

The IEA projects global data-centre electricity demand near 945 TWh in 2030 in its central case.[10] Lawrence Berkeley National Laboratory places U.S. data-centre demand in a wide range of 521–843 TWh by 2030, with a 649 TWh reference case.[11] Read as: the forecast uncertainty is large, but every plausible path makes power, grid connections, cooling, land, and finance material complements to the digital service.

Electricity is not merely an environmental externality attached to AI after production. It enters the production process directly. So do substations, transmission queues, water, memory bandwidth, network latency, and the utilization rate of specialized equipment. Ten accelerators waiting behind a delayed grid connection do not provide ten units of capital service. Nameplate capacity is not productive capacity until the full complement is available.

This physical base produces a two-capital geography. Regions with cheap, reliable power and permissive construction can attract physical AI capital. Regions and firms with strong skills, institutions, process data, and organizational capital can capture application value and TFP. Cloud services connect both places, but data sovereignty, regulation, latency, and grid congestion keep part of the geography local. Capital deepening and productivity rents need not accrue in the same place.

Falling inference prices also create rebound. Lower unit cost expands feasible use cases, lengthens context and reasoning, increases candidate generation, and allows agents to call tools and one another. Total compute and electricity demand can rise while the price per token falls. Efficiency is not the same as conservation; under elastic demand, it becomes the mechanism through which consumption expands.

The price index must therefore hold a quality-adjusted task constant. A token price can collapse while the number of tokens required for an accepted outcome increases. Conversely, a more expensive model can lower total outcome cost by reducing review, retries, or error loss. The correct deflator follows the service. A string-length index follows the meter.

Section 08

The capital race and the Golden Rule

Phelps’s Golden Rule, applied within the Solow model, asks which steady-state capital stock maximizes consumption rather than output.[25]Applied to AI, the social condition compares the marginal product of additional AI capital with depreciation, dilution, opportunity cost, and external cost. Grid congestion, carbon, water, specialized-asset fire-sale risk, and safety costs belong in the condition when markets do not price them. The private race for capacity and the socially efficient stock are not identical by definition.

Social Golden Rule condition for AI capital
(6)
MPKAI=n+g+δAI+MECAI\operatorname{MPK}_{AI}=n+g+\delta_{AI}+\operatorname{MEC}_{AI}

MEC denotes marginal external cost: grid congestion, environmental load, financial fragility, and other unpriced effects.

Wachter and Wachter start from observed hyperscaler investment and ask what beliefs about future productivity can rationalize it.[12] Their calibrated rare-boom model permits a remarkably wide range of macro outcomes: additional cumulative GDP growth by 2030 spans roughly 5 to 58 percentage points, while the AI sector’s share spans about 8% to 39%. Read as: current capital expenditure embeds weight on very large productivity states, but its arithmetic does not identify which state will occur.

Rungcharoenkitkul’s BIS model supplies the strategic counterweight.[13] When firms compete for a few dominant positions, each invests partly to win and does not internalize the duplication created by the race. Investment reaches around 1.5 times the efficient level in the conservative calibration and approaches three times when demand is less elastic. Read as: a real, consequential technology can still be overbuilt because the private contest and the social return solve different optimization problems.

Debt, circular equity stakes, and specialized assets add a financial layer. If realized demand or productivity disappoints, equipment with low alternative value can enter fire sales, while losses travel through connected balance sheets. None of this proves a bubble. It proves that “firms are investing” is evidence of belief, strategic necessity, financing conditions, and risk tolerance — not proof of realized productivity.

The railway analogy is useful only when anchored. Railways changed production and geography, and railway manias still destroyed capital. The implication is not that data centres are railways. The implication is that technical centrality and investment efficiency are separate propositions, and history permits one to be true while the other is false.

Section 09

Adoption is a ladder, not a switch

The adoption debate has acquired a peculiar empirical quality: several surveys appear correct and appear to disagree. The U.S. Census Bureau reports recent business use around 17–20%.[14] Eurostat reports AI use at about one in five EU enterprises with at least ten employees.[15] The ECB, asking a broad set of euro-area firms about several forms of AI, finds some use around 70%, while only 7% report significant use.[16] Read as: these are different points on a deployment process, not competing estimates of one homogeneous variable.

Figure · observed estimates + analytical ladder

Adoption is not one number.

Surveys disagree because they often measure different rungs. Exposure, active use, significant deployment, and production-recipe change are related variables, not synonyms.

  1. 01

    Exposure

    An employee opens a tool or uses an embedded feature.

    Broad and cheap to reach

  2. 02

    Experiment

    A team tests a bounded use case with low-risk data.

    Capability, not production

  3. 03

    Active use

    AI appears in current business operations with an owner and budget.

    17–20% of U.S. firms

  4. 04

    Significant use

    AI operates across meaningful functions and work volume.

    7% of euro-area firms

  5. 05

    Recipe change

    Roles, controls, products, or sequencing change durably.

    The early TFP indicator

U.S. current-use range: Census BTOS, May 2026. Euro-area significant use: ECB SAFE modules, published July 2026. The ECB also reports around 70% “some use,” illustrating why definitions must travel with the statistic. The estimates cover different populations and are not directly comparable.

At the bottom sits exposure: an employee opens a public tool or uses an embedded feature. Next comes experimentation, often with synthetic data and low-risk work. Active use places the system inside current operations. Functional deployment gives it permissions, support, ownership, and a budget. Workflow penetration measures the share of eligible work affected. At the top sits recipe change, where roles, steps, controls, sequencing, or products change because the system is trusted at scale. Access is sufficient for the first rung. Complementary capital governs the upper rungs.

Firm size exposes the mechanism. Census reports current use at 37% among firms with at least 250 employees, well above the economy-wide rate.[14] Eurostat reports 55% among large enterprises and 17% among small ones.[15] Read as: larger firms do not merely purchase more licences; they are more likely to possess process volume, data infrastructure, compliance capability, specialists, and managerial slack that convert access into use.

This creates an AI version of conditional convergence. Falling API prices do not force firms toward one productivity level because their steady states differ with data quality, process modularity, skills, security constraints, management, energy prices, and organizational investment. A common model can diffuse rapidly while productive intensity diverges. Democratized access and concentrated returns can therefore occur at the same time.

Europe’s debate often describes the transatlantic gap as a model-access problem. Access matters, but the ECB firms identify skills, privacy, and incompatibility with existing systems as central barriers.[16] Those are deficiencies in human, legal, data, and organizational capital. Subsidizing another proof of concept can increase experimentation while leaving the production function unchanged.

The measurement implication is straightforward. Report four quantities separately: the share of firms with any use, the share of workers using AI, the share of work hours affected, and the share of eligible workflows redesigned. The first is diffusion. The second and third capture intensity. The fourth is the closest early indicator of a production-recipe change. Without the ladder, adoption statistics generate more heat than knowledge.

Section 10

Why serious estimates disagree

Acemoglu’s task-based model adds around a tenth of a percentage point to annual productivity growth.[17] Filippucci, Gal, and Schief permit nearly a full point in their high sectoral scenario.[18] Jones and Tonetti generate a much larger acceleration when AI automates successive weak links, including idea production.[8]Averaging these results into a consensus forecast would remove the information contained in the disagreement. They describe different chains of assumptions and, in some cases, different economies.

macro gain = exposure × capability × adoption × net task saving × system translation × sector weight × equilibrium adjustment

Start with the share of tasks exposed to the technology. Exposure is not capability, so apply a capability rate, i.e. the share of tasks the current system can perform at the required quality. Capability is not adoption, so specify an adoption path. Adoption is not cost saving, so subtract review, errors, and integration from the local saving. A task saving is not a sector gain, so account for the task’s weight and remaining complements. A sector gain is not an aggregate gain, so include sector size, demand substitution, factor mobility, and prices. Finally, a level gain is not a new trend, so state whether the mechanism repeats through ideas and successive automation.

The terms interact. Heavy verification lowers net saving and slows adoption. Falling prices can raise adoption and trigger rebound. Task simplification changes labour supply, matching, and wages. Capital constraints raise prices where exposure is highest. Market power determines whether a technical saving becomes a lower customer price, a higher profit margin, or an unmeasured quality increase. The product of the terms is not a mechanical spreadsheet.

Acemoglu’s relatively modest macro estimate is driven by a limited set of tasks that can be profitably automated and conservative task-level savings.[17] Filippucci, Gal, and Schief obtain up to 0.9 percentage points of additional annual TFP in their high scenario when adoption reaches 40% of exposed tasks, savings are stronger, and sectors and factors adjust.[18] Read as: neither number is an observation of realized aggregate AI productivity; each is a transparent conditional result.

Jones and Tonetti show why even broad automation can have a muted aggregate effect when unautomated tasks remain weak links, while their more aggressive endogenous-automation path can eventually break the historical growth pattern.[8] The key parameters are now visible: productive adoption depth, task-to-system attenuation, organizational investment, effective depreciation, and the elasticity among tasks. None can be inferred from a model benchmark alone.

The micro evidence is consistent with this layered account. Brynjolfsson, Li, and Raymond document sizeable customer-support gains, concentrated among less experienced workers.[19] The EIB firm study estimates a 4% labour-productivity level effect among European adopters, driven by capital deepening rather than short-run job losses.[20] Yet Yotzov and co-authors report that most surveyed businesses still see no realized productivity effect.[21] Read together, the studies do not say “AI works” or “AI does not work.” They identify local efficacy, firm adoption, and aggregate productivity as separate stages of diffusion.

The responsible presentation is therefore a scenario decomposition. State the exposure set, capability threshold, adoption path, verification burden, elasticity assumptions, and whether the output is a level or growth-rate effect. Then show which parameter drives the result. False precision begins when the assumption chain is hidden behind one headline number.

Section 11

What a company should measure

Most enterprise AI dashboards answer the wrong question with increasing precision. They report active users, prompts, tokens, licences, latency, and perhaps estimated employee time saved. These are useful operating signals. They are not a production account, because they stop before the customer, release, payment, clinical event, or operational outcome that gives the activity economic meaning.

A serious dashboard begins downstream with five observations: the verified outcome, the measured input, the migrated bottleneck, the value retained after a model change, and the production-recipe change. These observations are simple to state because the instrumentation required to establish them is not.

  1. First

    Verified outcome productivity

    Accepted outcome value divided by inference, infrastructure, integration, review, rework, and expected error loss.

  2. Second

    Task-to-system attenuation

    The downstream outcome gain divided by the local task gain. This separates writing code from shipping it.

  3. Third

    Orchestration friction

    Review, coordination, exception handling, and recovery time per successful outcome.

  4. Fourth

    Capital durability

    The share of value retained after a model, hardware, or architecture migration.

  5. Fifth — and most crucially

    Recipe change

    The share of workflows whose steps, roles, controls, or sequencing changed durably.

The fifth measure is difficult because it asks whether the organization has become a different producer rather than a faster consumer of model output. It is also the measure closest to TFP in the economically meaningful sense. A licence can deepen capital. A durable elimination of handoffs, superior allocation, a new product, or a changed decision architecture can alter the production recipe.

Every serious deployment should therefore create an analysis-ready event log. Record the workflow and task, model and version, tokens and cost, latency and retries, tool calls, evaluation result, human review minutes, rework, escalation, acceptance, downstream completion, realized value, and error loss. Without the downstream event, the system measures activity. Without human effort, it understates cost. Without version and task identifiers, it cannot estimate depreciation or substitution.

Causal design matters because adopters are not random: current surveys associate adoption with firm size, age, sector, and digital capability.[14][16]Randomized rollout, staggered adoption with matched controls, valid instruments, production-function estimation, and fixed-task measurement experiments answer different parts of the problem. Every study should report local task performance, end-to-end production, and measured economic value. A gain in the first and not the latter two is not a null; it identifies the weak link.

The executive dashboard can be concise because the underlying account is complete. Verified outcomes, total cost, migrated bottlenecks, retained value, changed recipes. Tokens remain useful telemetry, but they no longer impersonate productivity.

Section 12

The counterargument: a stronger AI future

There is a future in which much of this caution becomes obsolete. If AI can perform nearly all economically essential cognitive tasks, learn the remaining ones, improve the research process that improves AI, and operate with little human verification, weak links can disappear rather than migrate. Intelligence would become reproducible at the margin, idea production would accelerate, and the economy could leave a familiar Solow transition for an endogenous-growth regime.

The present framework does not deny that future. It states what evidence would indicate it. Task-level gains would translate nearly one-for-one into final output. Human review would approach zero without expected loss rising. Elasticities would converge upward across domains. Production systems would preserve reliability over long horizons. Physical, institutional, and trust constraints would themselves become automatable. The residual would be supported by identified mechanisms rather than timing alone.

The claim that tokens are not a factor would weaken if a stable, homogeneous token unit emerged whose quality and economic service were invariant across models, tasks, and time, and if decomposing it into capital and intermediates destroyed explanatory power. The weak-link account would weaken if local gains translated one-for-one through diverse workflows with negligible review and integration. The organizational-capital account would weaken if firms obtained persistent causal output gains without changes in data, skills, processes, evaluations, or decision rights.

Current evidence is far from those conditions, but a model should remain open to being falsified. This is the difference between scepticism and discipline. Scepticism can become another prior. Discipline specifies which observation would make the prior untenable.

Conclusion

The old model and the new machine

Solow’s great contribution was not a particular exponent. It was a way of separating accumulation from technical change and a refusal to call every increase in output progress. That discipline is exactly what the machine economy needs. It permits the technology to be consequential in several ways at once without mistaking one mechanism for another.

AI can deepen capital, augment labour, substitute within tasks, improve the productivity of machines, create products, accelerate ideas, and change the production recipe. It can also increase electricity demand, depreciate rapidly, concentrate rents, erode learning paths, migrate bottlenecks, and invite over-investment. The question is not whether AI is “really” capital or “really” technology. The question is which mechanism operates, at which layer, with which complement, and whether the evidence can distinguish it.

Tokens will keep multiplying. Their prices will keep changing. Dashboards will keep counting them because they are there to be counted. But one more token is not one more unit of intelligence, labour, or output. It is one more attempt inside a production system.

Productivity begins when the attempt survives the system — and changes it.

Read.

the mechanism

Measure.

the complete system

Change.

the production recipe

Explore the Economics Evolving modelContinue to Part II: The Bottleneck Has Moved

Linked bibliography

References

Research cut-off: 2 August 2026. Status labels distinguish provisional working papers and research notes from journal and official statistical sources.

  1. 1

    Solow, Robert M. (1956).

    A Contribution to the Theory of Economic Growth. Quarterly Journal of Economics, 70(1), 65–94.

  2. 2

    Bryan, Kevin A. & Gans, Joshua S. (2026).

    Training AI for When Humans Will Use It. NBER Working Paper 35490.

    Working paperSource
  3. 3

    U.S. Bureau of Economic Analysis & Bureau of Labor Statistics (2026).

    Integrated Industry-Level Production Account for the United States. 1997–2024 release.

  4. 4

    U.S. Bureau of Labor Statistics (2026).

    AI and the Rise of Software Investment. Monthly Labor Review.

  5. 5

    Brynjolfsson, Erik; Rock, Daniel & Syverson, Chad (2021).

    The Productivity J-Curve. American Economic Journal: Macroeconomics, 13(1).

  6. 6

    Boyle, Shane; Fernald, John & Li, Huiyu (2026).

    Higher Utilisation Explains the Recent Surge in Productivity Growth. CEPR / VoxEU.

    Research noteSource
  7. 7

    Demirer, Mert; Musolff, Leon & Yang, Liyuan (2026).

    Writing Code vs. Shipping Code. NBER Working Paper 35275.

    Working paperSource
  8. 8

    Jones, Charles I. & Tonetti, Christopher (2026).

    Past Automation and Future A.I.: How Weak Links Tame the Growth Explosion. Stanford / NBER working paper, version 0.5.

    Working paperSource
  9. 9

    Althoff, Lukas & Reichardt, Hugo (2026).

    Task-Specific Technical Change and Comparative Advantage. NBER Working Paper 35353.

    Working paperSource
  10. 10

    International Energy Agency (2025).

    Energy Demand from AI. Energy and AI.

  11. 11

    Lawrence Berkeley National Laboratory (2026).

    United States Data Center Energy Usage Report: 2025 Update. LBNL.

  12. 12

    Wachter, Jessica & Wachter, Jonathan (2026).

    What Investment Data Implies about the AI Transition. NBER Working Paper 35290.

    Working paperSource
  13. 13

    Rungcharoenkitkul, Phurichai (2026).

    The AI Investment Race. BIS Working Paper 1367.

    Working paperSource
  14. 14

    U.S. Census Bureau (2026).

    Large Firms With at Least 20 Employees Biggest AI Users. Business Trends and Outlook Survey.

  15. 15

    Eurostat (2026).

    Use of Artificial Intelligence in Enterprises. 2025 enterprise data, June 2026 update.

  16. 16

    Ferrando, Annalisa; Lamboglia, Sara; Rariga, Judit & Schmidt, Maurice (2026).

    Adoption and Investment in AI across the Euro Area. ECB Occasional Paper 395.

  17. 17

    Acemoglu, Daron (2024).

    The Simple Macroeconomics of AI. NBER Working Paper 32487.

    Working paperSource
  18. 18

    Filippucci, Francesco; Gal, Peter & Schief, Matthias (2026).

    Aggregate Productivity Gains from Artificial Intelligence: A Sectoral Perspective. AEA Papers and Proceedings, 116.

  19. 19

    Brynjolfsson, Erik; Li, Danielle & Raymond, Lindsey R. (2025).

    Generative AI at Work. Quarterly Journal of Economics, 140(2), 889–942.

  20. 20

    European Investment Bank (2026).

    AI Adoption, Productivity and Employment: Evidence from European Firms. EIB Working Paper 2026/02.

    Working paperSource
  21. 21

    Yotzov, Iliyan et al. (2026).

    Firm Data on AI. NBER Working Paper 34836.

    Working paperSource
  22. 22

    Caplin, Andrew (2026).

    Planning Capital and Discovery-Based Learning-by-Doing. NBER Working Paper 35349.

    Working paperSource
  23. 23

    Borri, Nicola; Liu, Yukun & Tsyvinski, Aleh (2026).

    AI Premium. NBER Working Paper 35451 / arXiv.

    Working paperSource
  24. 24

    Acemoglu, Daron & Restrepo, Pascual (2018).

    The Race between Man and Machine. American Economic Review, 108(6), 1488–1542.

  25. 25

    Phelps, Edmund S. (1961).

    The Golden Rule of Accumulation: A Fable for Growthmen. American Economic Review, 51(4), 638–643.

All essaysCapital · services · outcomes