ThinkingResearch manuscript - second revision

Dependence architecture - welfare exposure - model ambiguity

Systemic Catastrophe under Architecture and Ambiguity

Component reliability does not determine systemic safety; dependence architecture does.

The paper derives joint failure from a latent-factor architecture, then asks whether its safety rate outruns both welfare exposure and the ambiguity surrounding the evidence.

AuthorMichael Schymura
VersionSecond revision - August 2026
Length39 pages
StatusUnpublished research manuscript

The central result

Three rates, one welfare boundary

A small probability attached to a large loss is not yet a model of catastrophic risk. It omits the rate at which probability falls, the rate at which consequence grows, the dependence architecture that generated the event, and the evidence that makes the probability defensible. Near a welfare boundary, those omissions decide the result.

min{Is(α),R}g(γ1)\min\{I_s(\alpha),R\}\gtrless g(\gamma-1)

Architecture generates the nominal systemic-safety exponent I. A defensible ambiguity set contracts at rate R. Welfare exposure grows at rate g, magnified under CRRA utility by curvature gamma. Robust tail pressure falls only when usable safety progress - the smaller of architecture and evidence - outruns exposure growth in utility units.

01
Section 3Analytical

One vanishing probability, three welfare regimes

Probability can fall while tail pressure vanishes, remains finite, or becomes dominant. The joint path, not the probability alone, decides the regime.

Three panels show catastrophic tail pressure tending to zero, a finite positive constant, or infinity as catastrophe probability falls.

Figure summary - Three panels show catastrophic tail pressure tending to zero, a finite positive constant, or infinity as catastrophe probability falls.

Interactive phase boundary

Which rate is actually winning?

Architecture supplies a nominal safety rate. Ambiguity limits the rate that can be defended. Welfare exposure sets the burden both must outrun.

Current regime

Robust tail pressure rises

Nominal architecture0.270
Evidence / ambiguity0.120
Usable safety = min(I, R)0.120
Welfare exposure = g(gamma - 1)0.160
Usable safety0.120
Exposure burden0.160
Safety margin-0.040

Exposure growth exceeds the defensible safety rate. A falling nominal failure probability is not enough to reduce robust tail pressure.

Section 01

The contribution is composition, not a new tail law

Three mature literatures meet in this problem. Catastrophe economics studies utility near a lower welfare boundary. Common-factor models derive joint loss from conditional dependence. Robust decision theory disciplines ambiguity and misspecification. The paper does not relabel any of them as new. Its contribution is to order their rates when architecture becomes a design choice.

ObjectEstablished comparatorWhat this paper adds
Catastrophic welfareBuchholz-Schymura susceptibility; regular variation; domain restrictionsA pathwise finite middle regime and an explicit rate comparison
Systemic probabilityConditional Bernoulli mixtures and common-factor loss limitsArchitecture as a safety choice with a least-cost route to joint failure
AmbiguityMaxmin, smooth ambiguity, multiplier preferences, KL neighborhoodsThe contraction rate R as a binding limit on usable safety progress
AI governanceGrowth, adoption, learning, irreversibility, and externality modelsA single boundary linking deployment architecture, exposure, evidence, and legacy

This also corrects the interpretation of the 2012 paper with Wolfgang Buchholz. Lower-unbounded utility establishes susceptibility: a tyrannical probability-severity path exists. It does not establish that every scientifically admissible path is tyrannical. Once the path is specified, tail pressure may vanish, converge to a finite positive value, or diverge.

02
Section 3.4Analytical

The CRRA boundary is a rate comparison

For the illustrative path c = p^b, the critical line b = 1/(gamma - 1) separates probability-dominated and severity-dominated regions.

A phase diagram plots the CRRA curvature parameter against the severity exponent and marks the boundary between negligible and dominant tail pressure.

Figure summary - A phase diagram plots the CRRA curvature parameter against the severity exponent and marks the boundary between negligible and dominant tail pressure.

Section 02

Architecture determines the least unlikely route to joint failure

Replication does not imply diversification. Systems can share a foundation model, cloud provider, identity plane, training corpus, orchestration layer, monitoring stack, or theory of control. Conditional on an architecture-wide factor, local failures may be independent; unconditionally, the system remains dependent.

Is(α)=infz{Js(z)+dα(qs(z))}I_s(\alpha)=\inf_z\left\{J_s(z)+d_\alpha\left(q_s(z)\right)\right\}

The infimum has a concrete meaning. Systemic failure can arrive through an unusually adverse common state, through an unusual concentration of local failures, or through a cheaper combination of both. Hardening therefore means more than improving average endpoint accuracy. It can require weaker common-factor loadings, independent fallbacks, recoverable control planes, and a less concentrated factor distribution.

Baseline architecture

Asymptotic0.1185
Finite N = 6400.1231

Hardened architecture

Asymptotic0.2700
Finite N = 6400.2749

The finite-system bridge verifies convergence for the stated Gaussian-logistic primitive. It is not a calibration to deployed AI systems.

05
Sections 6.3 and 6.7Structural

Diversification stops at the common-mode floor

Independent local failures diversify. Shared models, providers, identity planes, data, and oversight can leave a system-wide floor that component scores do not reveal.

A system diagram contrasts independent component failures with a shared common-mode dependency that can affect every component.

Figure summary - A system diagram contrasts independent component failures with a shared common-mode dependency that can affect every component.

Section 03

A safer design can outrun the evidence supporting it

Model ambiguity changes the usable rate. If the nominal event probability decays at exponent I while a binary KL neighborhood contracts at exponent R, the worst defensible probability decays at the smaller exponent. Engineering can improve faster than confidence in the engineering.

Ambiguity rate R0.04Hardened robust exponent0.0478
Ambiguity rate R0.12Hardened robust exponent0.1271
Ambiguity rate R0.30Hardened robust exponent0.2749

With a nominal hardened exponent near 0.27, an ambiguity rate of 0.04 leaves a robust exponent near 0.048. The design is materially safer inside the model. The admissible model set has not contracted at the same speed. A safety case should report both facts.

03
Section 6.1Analytical

Safety progress must outrun welfare exposure

Residual hazard can decline and tail pressure can still rise when welfare exposure deepens faster in utility units.

A phase field compares hazard-decay and welfare-exposure rates, with a diagonal boundary separating falling from rising tail pressure.

Figure summary - A phase field compares hazard-decay and welfare-exposure rates, with a diagonal boundary separating falling from rising tail pressure.

Section 04

Rollback stops new exposure; it does not rewrite history

Deployment changes the state from which later choices are made. Weights can be copied, access can persist, workflows can become dependent, and a latent trigger can exist before its consequence manifests. The paper therefore tracks capability, safety capital, and legacy as separate states.

Lresidual=μρ+μ+κtDL_{\mathrm{residual}}=\frac{\mu}{\rho+\mu+\kappa}\,\ell_tD

Here, ell is the posterior probability that a trigger is already latent, mu its manifestation rate, kappa the remediation rate, and D the continuation loss. Even when new deployment stops, residual expected loss remains positive whenever inherited latent exposure is positive.

This is why “deploy to learn” and “wait until we know” are not opposing theories. Controlled deployment has option value when information is available only through use and containment is credible. Waiting dominates when evidence can arrive without deployment or when a pilot can create irreversible legacy. The label “sandbox” does not settle which case applies.

Section 05

Routine success cannot cheaply validate an extreme tail

Under the favorable independent and stationary benchmark, zero observed catastrophes produce a one-sided upper bound. Certifying a probability below one in a million at 95 percent confidence requires almost three million zero-failure trials. Frontier AI violates stationarity, independence, and stable sampling almost by construction.

Elog(1/δ)ΔτE\geq\frac{\log(1/\delta)\,\Delta}{\tau}

The burden rises with the utility gap. As the welfare placed at risk grows, the acceptable probability shrinks and the evidence requirement grows with it. Pre-regime exposure gives no finite frequentist guarantee for a post-regime hazard after a material model, access, or deployment change.

04
Section 6.2Benchmark

Extreme reliability claims demand extreme evidence

Even the favorable independent and stationary benchmark requires nearly three million zero-failure trials to certify a probability below one in a million at 95 percent confidence.

A logarithmic chart shows the number of zero-failure trials required for increasingly small upper probability bounds.

Figure summary - A logarithmic chart shows the number of zero-failure trials required for increasingly small upper probability bounds.

Section 06

The code verifies the model, not the world

The companion simulation tracks capability, deployment, safety capital, legacy, and posterior hazard. It checks identities, compares direct and importance-sampling estimators, stresses structural parameters, and bridges finite systems to the asymptotic architecture theorem. Its fourth validation layer is a refusal: none of those checks counts as empirical validation of extinction risk.

Laissez-faire

9.20%

SE 0.409 pp

Staged release

1.88%

SE 0.192 pp

Safety first

0.62%

SE 0.111 pp

Adaptive governance

2.82%

SE 0.234 pp

Diversified stack

2.06%

SE 0.201 pp

Illustrative 40-period scenario outputs from 5,000 paths per policy. They compare stated parameterizations; they are not forecasts or empirical estimates.

Section 07

Replace one synthetic probability with an estimand map

The empirical program begins with mechanisms that can be measured, bounded, or falsified. A precise p(doom) that conceals capability, exposure, dependence, recovery, and welfare assumptions may contain less decision-relevant information than broad but transparent bounds.

ModuleCandidate evidenceCan identifyCannot establish alone
Normal benefitTask experiments, firm productivity, wages, pricesLocal causal effectsAggregate welfare or transformative growth
Capability and accessAutonomy horizons, tool evaluations, effective computeCapability in tested domainsCardinal catastrophe severity
Residual hazardPrespecified severe failures and exposure denominatorsRates on sampled regimesFrontier, adversarial, or regime-change hazard
Exposure and recoveryCritical-function adoption, rollback time, recoverabilitySystem exposure and resilienceThe sign of concentration for global risk
Dependence and propagationShared dependencies and empirical offspring matricesCommon-mode floors and local cascadesThe global tail without severity and saturation
WelfareConsumption, mortality, agency, population, continuation valueDeclared welfare componentsA unique value of extinction

Prospective identification requires a prespecified task distribution, compute budget, model version, access mode, safeguards, severity rule, and stopping boundary. Incident registries need common denominators such as model-hours, autonomous-action-hours, and critical-function exposure. Expert probabilities remain beliefs, not repeated-event frequencies.

Section 08

Extinction, agency, and consumption are not one variable

The original scalar catastrophe is useful for fixed-person consumption risk. It is not a complete representation of mass mortality, permanent institutional disempowerment, or extinction. Adding a constant to individual utility leaves fixed-population choices unchanged; under variable population it can reverse rankings. Ordinary risk aversion therefore cannot silently become the value of existence.

Δ=ΔC+ΔM+ΔA+ΔF\Delta=\Delta_C+\Delta_M+\Delta_A+\Delta_F

Consumption

Ordinary material loss and economic collapse

Mortality

Deaths, morbidity, and population change

Agency

Persistent loss of institutional and political control

Continuation

Foregone welfare of future recipients

The decomposition is valid only after mutually exclusive counterfactual increments are defined. Otherwise it double counts. The paper does not select a population ethic.

Section 09

Governance should target the rate that is binding

The model does not prove that one instrument dominates. It shows why bonds, capital requirements, staged licensing, evaluations, incident disclosure, access restrictions, interoperability, and recovery exercises act on different margins.

  1. 01

    Exposure-adjust the safety claim

    Pair incident rates with critical-function exposure, autonomy, tool access, recoverability, and shared dependencies.

  2. 02

    Treat staged release as an experiment

    A pilot earns its name only when it changes future decisions while containing external hazard and persistent legacy.

  3. 03

    Regulate common modes, not model counts

    Diversification must reduce the shared floor and cascade reproduction rate, not merely add nominal providers.

  4. 04

    Audit ambiguity separately

    A safety case should report architectural improvement and the evidence that contracts the defensible model set.

  5. 05

    Match instruments to margins

    Bonds, capital requirements, licensing, evaluations, disclosure, and access controls target different terms in the model.

For a CIO, the immediate inventory is a dependency graph across providers, identity systems, data stores, orchestration, monitoring, human fallback, and critical processes. For a regulator, the test is whether experimentation produces information without exporting the tail. For a board, capability progress belongs beside exposure, recoverability, concentration, and legacy.

Formal result directory

The scaffolding behind the boundary

Theorem 3.2

General pathwise limit

Tn=pnΔnT_n=p_n\Delta_n

The certainty equivalent converges exactly when tail pressure converges; the limit can be negligible, finite, or dominant.

Theorem 3.4

Regular-variation boundary

T(c)=cβρLp(1/c)LΔ(1/c)T(c)=c^{\beta-\rho}L_p(1/c)L_\Delta(1/c)

Probability decay wins when beta exceeds rho; severity wins when rho exceeds beta; lower-order terms decide equality.

Theorem 4.4

Ambiguity-rate erosion

Irobust=min{Is,R}I_{\mathrm{robust}}=\min\{I_s,R\}

A slowly contracting ambiguity set can erase the rate gain delivered by a safer nominal architecture.

Proposition 5.5

Rollback leaves a legacy floor

Lresidual=μρ+μ+κtDL_{\mathrm{residual}}=\frac{\mu}{\rho+\mu+\kappa}\ell_tD

Stopping new deployment prevents new triggers but cannot remove risk inherited from a latent trigger.

Theorem 6.7

Endogenous architecture rate

Is(α)=infz{Js(z)+dα(qs(z))}I_s(\alpha)=\inf_z\{J_s(z)+d_\alpha(q_s(z))\}

Systemic failure follows the least unlikely combination of an adverse common state and concentrated local failures.

Theorem 6.9

Architecture-welfare-ambiguity boundary

min{Is(α),R}g(γ1)\min\{I_s(\alpha),R\}\gtrless g(\gamma-1)

Robust tail pressure falls only when usable safety progress outruns welfare-exposure growth in utility units.

Theorem 6.10

Cascade threshold

ρ(M)1\rho(M)\leq1

In the classical local branching approximation, cascades die out almost surely only when the offspring matrix is subcritical.

Theorem 6.12

Symmetric overdeployment

Ωai=(1λi)[PiΔ+PΔai]\Omega_a^i=(1-\lambda_i)[P_i\Delta+P\Delta_{a_i}]

Firms choose more capability than the social planner when they internalize only part of continuation loss.

Research manuscript

Read the complete argument, proofs, and bibliography.

The website dossier is a guided map. The PDF remains the authoritative second revision, including assumptions, proofs, appendices, simulation parameterization, and the full reference list.

Version
Second revision
Length
39 pages
Status
Unpublished research manuscript
Research cutoff
9 August 2026
File
PDF - 1.30 MB
Integrity
89e4b27c4279b2d9

Selected intellectual lineage

References

The paper contains the full bibliography and distinguishes peer-reviewed work, working papers, institutional reports, and theorem comparators.

  1. 01

    Buchholz & Schymura (2012). Expected Utility Theory and the Tyranny of Catastrophic Risks

  2. 02

    Weitzman (2009). On Modeling and Interpreting the Economics of Catastrophic Climate Change

  3. 03

    Martin & Pindyck (2015). Averting Catastrophes: The Strange Economics of Scylla and Charybdis

  4. 04

    Hansen & Sargent (2022). Structured Ambiguity and Model Misspecification

  5. 05

    Jones (2024). The AI Dilemma: Growth versus Existential Risk

  6. 06

    Acemoglu & Lensman (2024). Regulating Transformative Technologies

  7. 07

    Gans (2025). How Learning about Harms Impacts the Optimal Rate of Artificial Intelligence Adoption

  8. 08

    Liski & Salanie (2026). Catastrophes, Delays, and Learning

  9. 09

    Bengio et al. (2026). International AI Safety Report 2026