Skip to main content
study
Research Dossier2026

Current research manuscript on AI, systemic risk, and decision-making under deep uncertainty.

When does a falling failure probability still increase catastrophic risk?

When the defensible systemic-safety rate fails to outrun welfare exposure. The binding boundary compares dependence architecture, ambiguity contraction, and utility curvature rather than probability alone.

AgenticMethods
24 minadvanced
Key Numbers
min{I, R}Usable Safety RateClick to scroll ↓
g(γ−1)CRRA Welfare BurdenClick to scroll ↓
≈3.0MZero-Failure TrialsClick to scroll ↓
39 pp.Complete ManuscriptClick to scroll ↓

Authoritative source

Second revision, complete proofs and appendices

This dossier isolates the formal mechanism and its identification limits. The 39-page manuscript contains the complete proofs, simulation specification, robustness checks, and bibliography.

Formal object

Let zz denote the normal-state consumption equivalent, cnc_n the catastrophic-state consumption equivalent, and pnp_n the probability of that state along a sequence of systems indexed by scale nn. Define the utility gap and catastrophic tail pressure as

Δn=u(z)u(cn),Tn=pnΔn.\Delta_n = u(z)-u(c_n), \qquad T_n=p_n\Delta_n.

Expected utility can then be written without approximation:

E[u(Cn)]=u(z)Tn,CEn=u1 ⁣(u(z)Tn).\mathbb E[u(C_n)] = u(z)-T_n, \qquad CE_n=u^{-1}\!\left(u(z)-T_n\right).

This decomposition separates two objects that are often compressed into a single probability: the event rate pnp_n and the welfare exposure Δn\Delta_n. A falling pnp_n is sufficient for declining risk only when the product falls. The certainty equivalent inherits three pathwise regimes:

Limit of TnT_nWelfare regimeCertainty-equivalent implication
Tn0T_n\to0Probability dominatesCEnzCE_n\to z
Tnτ(0,)T_n\to\tau\in(0,\infty)Finite middle regimeCEnu1(u(z)τ)CE_n\to u^{-1}(u(z)-\tau)
TnT_n\to\inftySeverity dominatesCEnCE_n approaches the welfare boundary

The middle regime matters. It rules out the false dichotomy in which a vanishing probability must become either irrelevant or infinitely dominant.

Limit classification

Suppose probability and utility loss are regularly varying as catastrophic consumption approaches zero:

p(c)=cβLp(1/c),Δ(c)=cρLΔ(1/c),p(c)=c^{\beta}L_p(1/c), \qquad \Delta(c)=c^{-\rho}L_\Delta(1/c),

where LpL_p and LΔL_\Delta are slowly varying. Tail pressure is therefore

T(c)=cβρLp(1/c)LΔ(1/c).T(c)=c^{\beta-\rho}L_p(1/c)L_\Delta(1/c).

The leading exponents classify the limit. Probability decay wins when β>ρ\beta>\rho; welfare severity wins when ρ>β\rho>\beta; slowly varying terms decide the boundary β=ρ\beta=\rho. Under CRRA utility with curvature γ>1\gamma>1, Δ(c)\Delta(c) is asymptotically proportional to c1γc^{1-\gamma}. For the path cn=exp(gn)c_n=\exp(-gn), the welfare burden grows at exponent g(γ1)g(\gamma-1).

The CRRA phase boundary separates probability-dominated and severity-dominated paths.

Architecture exponent

Component reliability is not a systemic model. Let ZnZ_n be a common state generated by architecture ss, and let local components fail conditionally with probability qs(Zn)q_s(Z_n). For a systemic threshold α\alpha, the probability that at least αn\alpha n components fail depends on both an adverse common state and conditional concentration.

If ZnZ_n satisfies a large-deviation principle with rate function JsJ_s, the systemic-safety exponent is

Is(α)=infz{Js(z)+dα ⁣(qs(z))},I_s(\alpha)=\inf_z\left\{J_s(z)+d_\alpha\!\left(q_s(z)\right)\right\},

with binary relative entropy

dα(q)=αlogαq+(1α)log1α1q.d_\alpha(q)= \alpha\log\frac{\alpha}{q} +(1-\alpha)\log\frac{1-\alpha}{1-q}.

The infimum identifies the least unlikely route to joint failure: an unusually adverse shared state, an unusual concentration of local failures, or a cheaper combination of both. Shared foundation models, cloud control planes, identity systems, data pipelines, and monitoring assumptions enter through JsJ_s and qsq_s; adding nominally independent endpoints does not remove those common modes.

Independent local failures diversify, while common-mode dependence creates a systemic floor.

Ambiguity envelope

Assume the nominal systemic probability behaves as Pn,sexp[nIs(α)]P_{n,s}\asymp\exp[-nI_s(\alpha)]. A robust decision-maker does not know that nominal law exactly and evaluates probabilities over an ambiguity set that contracts with evidence. If the defensible set contracts at exponent RR, the worst-case probability has usable exponent

Irobust=min{Is(α),R}.I_{\mathrm{robust}}=\min\{I_s(\alpha),R\}.

This minimum is the technical reason a safer design can improve faster than the evidence supporting it. Architecture governs the nominal rate. Identification and validation govern how quickly nearby models can be rejected. The combined phase boundary is

min{Is(α),R}g(γ1).\boxed{\min\{I_s(\alpha),R\}\gtrless g(\gamma-1)}.

Robust tail pressure falls only when the smaller of architecture progress and ambiguity contraction exceeds welfare-exposure growth in utility units. Equality is not a knife-edge answer: finite-system corrections and lower-order terms become decisive.

Finite-system check

The asymptotic result must be connected back to systems of finite size. For the paper's Gaussian-logistic primitive at n=640n=640, the direct finite exponent is close to the asymptotic prediction:

ArchitectureAsymptotic exponentFinite exponent, n=640n=640Absolute gap
Baseline0.11850.11850.12310.12310.00460.0046
Hardened0.27000.27000.27490.27490.00490.0049

This is a numerical implementation check, not an empirical estimate of catastrophic AI risk. It establishes that the stated finite primitive approaches the theorem at the reported scale. It does not establish that deployed systems follow the primitive.

Identification boundary

Even the favorable independent and stationary benchmark imposes a severe validation burden. With zero observed failures in EE trials, a one-sided (1δ)(1-\delta) upper confidence bound is

pˉ(E,δ)=1δ1/E.\bar p(E,\delta)=1-\delta^{1/E}.

To certify pτp\le\tau, the required exposure is approximately

Elog(1/δ)τ.E\ge \frac{\log(1/\delta)}{\tau}.

At τ=106\tau=10^{-6} and δ=0.05\delta=0.05, this is approximately 2.9962.996 million zero-failure trials. Frontier systems weaken every convenient assumption behind that number: model versions change, exposure is selected, outcomes are dependent, and deployment may move the operating regime. Routine success therefore identifies routine reliability, not automatically an extreme tail.

The empirical program should estimate separate margins: capability and access; critical-function exposure; incident denominators; recovery and rollback time; common dependencies; cascade reproduction; and ambiguity contraction. A single synthetic probability conceals which margin produced the result and which observation could falsify it.

Welfare boundary

Consumption collapse, mortality, loss of agency, and extinction cannot be folded into CRRA curvature without an additional ethical argument. A transparent accounting identity is

Δ=ΔC+ΔM+ΔA+ΔF,\Delta=\Delta_C+\Delta_M+\Delta_A+\Delta_F,

where the increments represent consumption, mortality, agency, and continuation losses under mutually exclusive counterfactuals. The decomposition is useful only when those counterfactuals prevent double counting. The paper deliberately does not infer a population ethic from ordinary risk aversion.

Policy margins

The model does not imply one universal instrument. It maps instruments to rates. Architecture standards, independent fallbacks, interoperability, and recovery exercises target Is(α)I_s(\alpha). Prespecified evaluations, incident denominators, disclosure, and adversarial testing target RR. Staged access, reversible deployment, and critical-function limits target gg. Liability, bonds, capital requirements, and insurance alter private incentives when firms internalize only part of continuation loss.

The practical test is therefore sharper than “Is the system safe?” A safety case should state which exponent it claims to improve, what evidence identifies that change, which common modes remain, and how much welfare exposure is accumulating while the evidence arrives.

Interactive phase boundary

Which rate is actually winning?

Architecture supplies a nominal safety rate. Ambiguity limits the rate that can be defended. Welfare exposure sets the burden both must outrun.

Current regime

Robust tail pressure rises

Nominal architecture0.270
Evidence / ambiguity0.120
Usable safety = min(I, R)0.120
Welfare exposure = g(gamma - 1)0.160
Usable safety0.120
Exposure burden0.160
Safety margin-0.040

Exposure growth exceeds the defensible safety rate. A falling nominal failure probability is not enough to reduce robust tail pressure.

Lab Panel

Switch to desktop view for the full Lab Panel experience with section-aware insights and callouts.