Authoritative source
Second revision, complete proofs and appendices
This dossier isolates the formal mechanism and its identification limits. The 39-page manuscript contains the complete proofs, simulation specification, robustness checks, and bibliography.
Formal object
Let denote the normal-state consumption equivalent, the catastrophic-state consumption equivalent, and the probability of that state along a sequence of systems indexed by scale . Define the utility gap and catastrophic tail pressure as
Expected utility can then be written without approximation:
This decomposition separates two objects that are often compressed into a single probability: the event rate and the welfare exposure . A falling is sufficient for declining risk only when the product falls. The certainty equivalent inherits three pathwise regimes:
| Limit of | Welfare regime | Certainty-equivalent implication |
|---|---|---|
| Probability dominates | ||
| Finite middle regime | ||
| Severity dominates | approaches the welfare boundary |
The middle regime matters. It rules out the false dichotomy in which a vanishing probability must become either irrelevant or infinitely dominant.
Limit classification
Suppose probability and utility loss are regularly varying as catastrophic consumption approaches zero:
where and are slowly varying. Tail pressure is therefore
The leading exponents classify the limit. Probability decay wins when ; welfare severity wins when ; slowly varying terms decide the boundary . Under CRRA utility with curvature , is asymptotically proportional to . For the path , the welfare burden grows at exponent .
Architecture exponent
Component reliability is not a systemic model. Let be a common state generated by architecture , and let local components fail conditionally with probability . For a systemic threshold , the probability that at least components fail depends on both an adverse common state and conditional concentration.
If satisfies a large-deviation principle with rate function , the systemic-safety exponent is
with binary relative entropy
The infimum identifies the least unlikely route to joint failure: an unusually adverse shared state, an unusual concentration of local failures, or a cheaper combination of both. Shared foundation models, cloud control planes, identity systems, data pipelines, and monitoring assumptions enter through and ; adding nominally independent endpoints does not remove those common modes.
Ambiguity envelope
Assume the nominal systemic probability behaves as . A robust decision-maker does not know that nominal law exactly and evaluates probabilities over an ambiguity set that contracts with evidence. If the defensible set contracts at exponent , the worst-case probability has usable exponent
This minimum is the technical reason a safer design can improve faster than the evidence supporting it. Architecture governs the nominal rate. Identification and validation govern how quickly nearby models can be rejected. The combined phase boundary is
Robust tail pressure falls only when the smaller of architecture progress and ambiguity contraction exceeds welfare-exposure growth in utility units. Equality is not a knife-edge answer: finite-system corrections and lower-order terms become decisive.
Finite-system check
The asymptotic result must be connected back to systems of finite size. For the paper's Gaussian-logistic primitive at , the direct finite exponent is close to the asymptotic prediction:
| Architecture | Asymptotic exponent | Finite exponent, | Absolute gap |
|---|---|---|---|
| Baseline | |||
| Hardened |
This is a numerical implementation check, not an empirical estimate of catastrophic AI risk. It establishes that the stated finite primitive approaches the theorem at the reported scale. It does not establish that deployed systems follow the primitive.
Identification boundary
Even the favorable independent and stationary benchmark imposes a severe validation burden. With zero observed failures in trials, a one-sided upper confidence bound is
To certify , the required exposure is approximately
At and , this is approximately million zero-failure trials. Frontier systems weaken every convenient assumption behind that number: model versions change, exposure is selected, outcomes are dependent, and deployment may move the operating regime. Routine success therefore identifies routine reliability, not automatically an extreme tail.
The empirical program should estimate separate margins: capability and access; critical-function exposure; incident denominators; recovery and rollback time; common dependencies; cascade reproduction; and ambiguity contraction. A single synthetic probability conceals which margin produced the result and which observation could falsify it.
Welfare boundary
Consumption collapse, mortality, loss of agency, and extinction cannot be folded into CRRA curvature without an additional ethical argument. A transparent accounting identity is
where the increments represent consumption, mortality, agency, and continuation losses under mutually exclusive counterfactuals. The decomposition is useful only when those counterfactuals prevent double counting. The paper deliberately does not infer a population ethic from ordinary risk aversion.
Policy margins
The model does not imply one universal instrument. It maps instruments to rates. Architecture standards, independent fallbacks, interoperability, and recovery exercises target . Prespecified evaluations, incident denominators, disclosure, and adversarial testing target . Staged access, reversible deployment, and critical-function limits target . Liability, bonds, capital requirements, and insurance alter private incentives when firms internalize only part of continuation loss.
The practical test is therefore sharper than “Is the system safe?” A safety case should state which exponent it claims to improve, what evidence identifies that change, which common modes remain, and how much welfare exposure is accumulating while the evidence arrives.
Interactive phase boundary
Which rate is actually winning?
Architecture supplies a nominal safety rate. Ambiguity limits the rate that can be defended. Welfare exposure sets the burden both must outrun.
Current regime
Robust tail pressure rises
Exposure growth exceeds the defensible safety rate. A falling nominal failure probability is not enough to reduce robust tail pressure.