0 cumulative citations
View corpus contextThe race for AGI can speed up as existential risk rises: when catastrophe costs enter both competitors' payoffs they cancel out, producing a 'suicide region' in which rational actors deploy despite negative social value; assigning liability or sharing prizes can restore welfare, while isolated warning-shot disasters fail to deter acceleration.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We analyze a continuous-time preemption game with shared catastrophic externalities. When the cost of catastrophe is embedded in both players' payoffs, the risk term cancels out in the equilibrium indifference condition. This creates a "suicide region" where competitive pressures force rational agents to deploy despite negative risk-adjusted net present values. We apply this framework to the race for artificial general intelligence (AGI). We show that this suicide region widens as the cost of systemic ruin grows: higher catastrophic risk does not deter the race but instead enlarges the set of conditions under which rational actors deploy despite negative social value. We characterize the resulting welfare distortion against a social planner's benchmark and demonstrate how two complementary mechanisms - private liability and prize-sharing - can close the suicide region. Private liability raises the cost of unsafe deployment while prize-sharing reduces the strategic imperative to deploy first. "Warning shots" (sub-existential disasters) will fail to deter AGI acceleration, as the winner-takes-all nature of the race remains intact.
Summary
Main Finding
When a preemptive race (winner-takes-all) is coupled with a catastrophic downside that is shared identically by all players, the catastrophic term cancels from the equilibrium indifference condition. As a result, increasing the size of systemic ruin does not deter acceleration — instead it enlarges a “suicide region” of parameter space in which rational actors deploy early (and even with negative risk‑adjusted NPV) to avoid being second. Two complementary mechanisms can eliminate this suicide region: (i) private liability that internalizes the catastrophe cost up to the social magnitude, and (ii) prize‑sharing (windfall clauses) that reduces the first‑mover payoff and hence the strategic imperative to preempt.
Key Points
- Cancellation effect: If the catastrophe cost (D) enters identically into all players’ payoffs, it drops out of the equilibrium indifference condition for timing. Therefore D does not enter the strategic timing decision once the race is underway.
- Suicide region: A region of model parameter space where competitive pressure forces deployment even though deployment is socially and privately risk‑adjusted negative. The suicide region widens monotonically with D.
- Mechanism design:
- Private liability: Imposing liability on leaders that internalizes the social cost (up to 2D in the two‑player case) can close the suicide region by raising the private cost of unsafe deployment.
- Prize‑sharing (S*): An optimal share of the winner’s windfall to others (e.g., windfall clauses) reduces the winner‑takes‑all character and eliminates the preemption incentive.
- Complementarity: Liability raises leader costs; prize sharing reduces strategic gain — together they address both sides of the payoff.
- Warning‑shot events (sub‑existential disasters) are unlikely to deter acceleration because the winner‑takes‑all dynamic still motivates rush to deploy.
- Generalizability: The mechanism applies beyond AGI — any race with a large shared catastrophic externality and large first‑mover advantage (e.g., autonomous weapons, weaponized space, risky synthetic biology) can exhibit the same dynamics.
- Limitations identified in the paper: enforcement of liability after an existential event is impossible, and some deterrence relies on ex‑ante enforceability; model is developed in a symmetric two‑player continuous‑time setting (extensions to heterogeneity/multi‑player noted as future work).
Data & Methods
- Nature of study: Theoretical, formal model (no empirical dataset).
- Modeling framework:
- Continuous‑time real‑options preemption game building on Weeds (2002), Huisman & Kort (2003), Grenadier (2002).
- Two symmetric players, winner‑takes‑all payoff structure, irreversible deployment decision.
- Key primitives and notation (as used/introduced in the paper): V (value of achieving AGI), I (sunk cost of deployment), σ (volatility; state‑dependent and increasing with V to capture recursive self‑improvement), D (monetized disutility of systemic ruin / existential tail), S (second‑mover share or prize‑sharing parameter).
- Novel ingredients:
- Endogenous ruin parameter D correlated with deployment velocity (captures alignment failure risk increasing with speed).
- Shared catastrophic externality — D enters every player's payoff identically.
- Knightian uncertainty over misalignment tails considered qualitatively when discussing liability deterrence.
- Analytical approach:
- Derive equilibrium indifference (threshold) conditions for waiting vs. deploying in continuous time; show algebraic cancellation of D across players’ payoffs.
- Characterize the suicide region via comparative statics (how thresholds move with D, σ, I, etc.).
- Construct social planner benchmark V_social* and compute welfare gap between decentralized equilibrium and planner optimum; show gap widens with D.
- Solve for boundary conditions that eliminate the suicide region: private liability magnitude (≥ 2D in the two‑player model) and optimal prize‑sharing fraction S* that removes winner‑take‑all incentives.
- Robustness/qualitative considerations:
- Discuss how ex‑ante liability works under Knightian uncertainty (cannot reliably discount existential tail), and how liability/verification/prize‑sharing interact during development vs. “breakout” stage.
Implications for AI Economics
- Rethinks deterrence intuition: Greater systemic/existential risk does not necessarily create restraint in competitive preemption races. Policy prescriptions based on “mutually assured catastrophe will restrain actors” are unreliable when catastrophe is shared.
- Policy levers:
- Private liability (civil/criminal penalties, enforceable compensation regimes) can be effective ex‑ante if designed to internalize social cost; requires credible international enforcement and pre‑deployment verification mechanisms.
- Prize‑sharing / windfall‑clause designs (contractual or regulatory) that redistribute a portion of the first mover’s gains reduce the strategic benefit of winning and thus the pressure to cut safety.
- Combined approaches are most robust: raising leader costs and lowering first‑mover gains simultaneously addresses the two forces driving the suicide region.
- Implementation challenges:
- Enforcement: Liability is weak ex‑post if an existential catastrophe occurs; its power rests on credible ex‑ante enforcement and verification.
- International coordination: The efficacy of both liability and prize sharing depends on cross‑jurisdictional agreements (export controls, treaties, multilateral windfall agreements).
- Asymmetric actors and many‑player races: The two‑player symmetric model is illustrative; heterogeneous capabilities, private vs state actors, and multi‑player dynamics could alter thresholds and require tailored instruments.
- Research and policy priorities suggested:
- Modeling extensions: multi‑player races, heterogeneous capabilities and costs, endogenous safety R&D investments, dynamic enforcement probabilities.
- Institutional design: international legal frameworks for pre‑deployment certification, enforceable windfall sharing, and mechanisms to credibly bind prolific actors.
- Empirical work: measure real‑world alignment tax tradeoffs, observe safety‑for‑speed behavior across labs/nations, and evaluate feasible liability/enforcement architectures.
- Broader message for AI economics: Strategic incentives can dominate risk magnitudes in high‑stakes technological races. Effective regulation must change payoff structures (not just raise awareness of risk) to prevent rational but socially catastrophic acceleration.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| When the cost of catastrophe is embedded in both players' payoffs, the risk term cancels out in the equilibrium indifference condition. Adoption Rate | null_result | dependence of the equilibrium indifference condition on the catastrophic risk term |
Reading fidelity
high
Study strength
medium
|
not reported
|
| This creates a 'suicide region' where competitive pressures force rational agents to deploy despite negative risk-adjusted net present values. Adoption Rate | positive | existence and extent of a parameter region in which agents deploy despite negative risk-adjusted NPV (the 'suicide region') |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The suicide region widens as the cost of systemic ruin grows: higher catastrophic risk does not deter the race but instead enlarges the set of conditions under which rational actors deploy despite negative social value. Adoption Rate | positive | size/measure of the suicide region as a function of the cost of systemic ruin |
Reading fidelity
high
Study strength
medium
|
not reported
|
| There is a welfare distortion relative to a social planner's benchmark resulting from the race dynamics (i.e., private equilibrium outcomes are inefficient compared to the social planner). Consumer Welfare | negative | welfare gap / distortion between decentralized equilibrium and social planner optimum |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Two complementary mechanisms—private liability and prize-sharing—can close the suicide region. Adoption Rate | negative | presence/size of the suicide region under policy interventions (private liability, prize-sharing) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Private liability raises the cost of unsafe deployment. Governance And Regulation | negative | private cost of unsafe deployment |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Prize-sharing reduces the strategic imperative to deploy first. Adoption Rate | negative | strategic incentive to deploy first (first-mover advantage) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| 'Warning shots' (sub-existential disasters) will fail to deter AGI acceleration, as the winner-takes-all nature of the race remains intact. Adoption Rate | null_result | effectiveness of sub-existential 'warning shot' disasters in deterring AGI acceleration |
Reading fidelity
high
Study strength
medium
|
not reported
|