0 cumulative citations
View corpus contextRacing firms can push AI development to existential thresholds: with perfect observability and trust the technology frontier is bounded and coordinated stopping is feasible, but opacity or distrust can make perpetual risky racing a self-enforcing equilibrium that leads to disaster.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We study R&D competition in the shadow of disaster: advancing the technology frontier raises the risk of permanently ending all firms' payoffs. Under perfect monitoring and common knowledge of rationality, the equilibrium frontier is bounded below by the optimal stopping time of a monopolist, and above by that of a representative firm that persistently but mistakenly believes its rival is about to stop. We then analyze how the frontier is shaped by transparency (speed of monitoring) and trust (belief in the rationality of rival firms).
Summary
Main Finding
Competition to scale dangerous technologies can push the technological frontier much farther than a monopolist (or a single optimistic firm) would choose, generating a risk of “racing to ruin.” Under perfect monitoring and common knowledge of rationality, the equilibrium frontier is always between two simple benchmarks: (i) a lower bound equal to the monopolist’s optimal stopping time (τ) and (ii) an upper bound equal to the stopping time of a representative firm that optimistically treats its rival’s technology as frozen (τR). Relaxing transparency (delayed detection) or trust (probability the rival is irrational) widens the set of possible outcomes and can generate equilibria in which firms race forever and disaster occurs with probability one. Transparency and trust interact non‑monotonically: faster detection and greater trust generally help coordination, but faster detection can, at intermediate trust levels, destroy some desirable equilibria before restoring them.
Key Points
- Model setup
- Two symmetric firms advance their own technology at unit speed until they irrevocably stop. The frontier is the maximum of the two technologies.
- Frontier scale raises an exogenous hazard λ(frontier) of a permanent industry collapse (flow payoffs drop to zero).
- Firms obtain flow profits π(own tech, rival tech) that increase in own tech and decrease in rival tech; firms discount at rate r.
- Representative-agent benchmark (τR)
- Define τR as the first time at which a firm that (mistakenly) treats its rival’s technology as frozen optimally prefers to stop.
- τR is an optimistic (upper) benchmark because it ignores the rival’s continued scaling.
- Monopolist benchmark (τ)
- Define τ as the stopping time of a monopolist (or when rival is at 0). This provides a lower bound on how far the frontier will be carried.
- Main equilibrium bounds (perfect transparency)
- Proposition 1 (Upper bound): In any subgame perfect equilibrium (SPE) with perfect monitoring, the realized frontier never exceeds τR.
- Proposition 2 (Lower bound): Across SPEs the frontier reaches at least τ.
- Hence the equilibrium frontier lies in [τ, τR]. In many parametric examples the bounds are close.
- Equilibrium structure (perfect transparency)
- There always exists a pure-strategy SPE (a “pure priority” equilibrium).
- Any realized outcome is either “spread” (firms stop at different times with frontier strictly inside [τ, τR]) or “bunched” (both stop simultaneously at τR).
- Comparative statics
- Increasing the marginal benefit of own scaling (Di log π on the diagonal) raises τR.
- Increasing the hazard λ lowers τR and τ.
- Imperfect transparency (delayed detection)
- Stopping is observed after an exponentially distributed delay with rate ω (higher ω = faster monitoring).
- If monitoring is very noisy (small ω), there can be equilibria where neither firm ever stops and disaster occurs with probability 1.
- When monitoring is sufficiently precise, all equilibria stop in finite time; the expected overshoot of the frontier beyond τR is at most O(1/ω).
- Trust / incomplete beliefs about rationality
- Introduce a small probability that the rival is a “crazy” type who never stops.
- Take a stationary environment where τR = 0 (at any state, a firm wants to stop if rival also does).
- Results partition prior belief in rationality into three zones:
- Low trust: every equilibrium races to ruin (disaster probability 1).
- Intermediate trust: both immediate stopping and racing-to-ruin can be equilibria.
- High trust: probability of two rational firms racing forever vanishes (specifically vanishes quadratically with the prior odds of rationality).
- Transparency is double-edged: increasing observation speed can first eliminate cooperative equilibria by making free-riding easier, then restore them when detection becomes fast enough.
- Relation to literature
- Connects dynamic R&D races, preemption, wars of attrition, and contest literature to questions of existential/transformative risk in AI and other scaling technologies.
Data & Methods
- Type of study: Theoretical / analytical game-theoretic model (continuous-time stochastic stopping game).
- Key modeling choices and assumptions
- Two symmetric firms; technology advances deterministically at unit speed until stopping; stopping is irreversible.
- Frontier hazard λ(t) is nondecreasing in frontier scale; flow payoffs π(t_i, t_j) are smooth, positive, increasing in own tech, decreasing in rival tech, with log-supermodularity and single-crossing properties (Assumption 1).
- Perfect monitoring case: firms observe each other’s stops immediately; analyze subgame perfect equilibria.
- Imperfect monitoring extension: detection lag exponential with rate ω; analyze how equilibria depend on ω.
- Trust extension: incomplete information with a small mass of “crazy” (never-stop) types; follow Kreps–Wilson–style signaling reasoning in continuous time.
- Analytical methods
- Construct and analyze a representative (optimistic) stopping problem to define τR.
- Use stopping-time arguments, single-crossing and log-supermodularity to derive upper and lower bounds on equilibrium frontiers.
- Backward-induction and deviation arguments for equilibrium characterization with perfect information.
- Study limit behavior and comparative statics; derive rates of overshoot O(1/ω) and vanishing probabilities in the trust extension.
- No empirical data or calibration in the paper; results are qualitative and robust across parametric examples.
Implications for AI Economics
- Strategic source of excessive risk: Even with risk-neutral, patient agents, competition can drive frontier scaling beyond socially or privately optimal points (relative to a monopolist), because firms face relative-payoff incentives and coordination problems.
- Role of transparency
- Faster, reliable monitoring of rivals’ stopping behavior generally helps coordination and limits runaway scaling, but the effect is non-monotone when beliefs about rivals’ rationality are uncertain: intermediate increases in monitoring speed can destabilize cooperative equilibria by facilitating opportunistic free-riding.
- Policy implication: design of monitoring systems and disclosure regimes should consider the interaction with beliefs about compliance and incentives to free-ride.
- Role of trust and institutions
- Trust (credible belief that rivals are rational and will stop) is crucial. Low trust makes catastrophic racing inevitable; raising trust (via commitments, verification, enforceable agreements, third-party certification) can sharply reduce the chance of ruin.
- Institutions that credibly change opponents’ type (e.g., verified commitments, binding international agreements, licensing or certification with enforceable penalties) can move the game from low- to high-trust regimes.
- Multiple equilibria and fragility
- Because equilibria can include both cooperative stopping and racing-to-ruin, small changes in monitoring, priors, or enforcement can flip outcomes. Policy should therefore aim both to shift incentives and to stabilize the cooperative equilibrium (e.g., by making stopping self-enforcing).
- Practical levers suggested by the model
- Improve monitoring speed and reliability (increase ω), while also ensuring that monitoring changes do not inadvertently enable free-riding absent sufficient trust.
- Increase transparency plus credible enforcement to transform unilateral stopping into a mutually observable, enforceable action.
- Create institutions that raise the prior odds that firms are “rational” (i.e., committed to safety), e.g., industry standards, binding multilateral pacts, or third-party audits.
- Consider altering payoff structure (e.g., taxes/subsidies or liability rules) to reduce marginal returns to scaling and thereby lower τR.
- Limitations to bear in mind
- Stylized two-firm model; real-world settings have many developers, heterogeneous capabilities, and richer strategy spaces (e.g., partial slowdown, safety investments).
- Hazard depends only on the leader’s frontier in the model; in practice, systemic risk could depend on other features (usage, diffusion, feedback loops).
- No explicit regulator or side-payments; adding enforceable contracts or centralized authority could change equilibrium logic.
- Results are qualitative; calibration to AI developer data would be needed to assess magnitudes and concrete policy thresholds.
Overall, the paper provides a clear, tractable framework showing how competition, imperfect monitoring, and distrust can jointly generate excessive scaling and existential risk — and identifies transparency and trust as central, interacting policy levers.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Under perfect monitoring and common knowledge of rationality, the equilibrium technology frontier is bounded above by the representative firm's stopping threshold τR. Ai Safety And Ethics | negative | Technology frontier reached before firms stop and disaster risk is eliminated |
Reading fidelity
high
Study strength
high
|
n=2
|
| Under perfect monitoring, every equilibrium advances the technology frontier to at least the stopping time of a monopolist, denoted τ. Ai Safety And Ethics | negative | Minimum technology frontier reached before stopping |
Reading fidelity
high
Study strength
high
|
n=2
|
| Under perfect transparency, the equilibrium technology frontier lies in the interval [τ, τR]. Ai Safety And Ethics | mixed | Equilibrium technology frontier before stopping |
Reading fidelity
high
Study strength
high
|
n=2
[τ, τR]
|
| There exists a pure-strategy subgame-perfect equilibrium in the baseline model. Market Structure | positive | Existence of a pure-strategy equilibrium |
Reading fidelity
high
Study strength
high
|
n=2
|
| In every perfect-transparency equilibrium, stopping outcomes are either spread, with the firms stopping at different times and the frontier strictly below τR, or bunched, with both firms stopping simultaneously at τR. Task Allocation | mixed | Distribution and timing of firms' stopping decisions |
Reading fidelity
high
Study strength
high
|
n=2
|
| When stopping decisions are observed with sufficiently noisy and slow monitoring, there are equilibria in which neither firm ever stops and disaster occurs with probability one. Ai Safety And Ethics | negative | Probability of disaster and whether firms ever stop |
Reading fidelity
high
Study strength
high
|
n=2
|
| When monitoring is sufficiently precise, every equilibrium stops in finite time, but the frontier can overshoot τR because each firm prefers to stop only after confirming that its rival has stopped. Ai Safety And Ethics | negative | Stopping time and technology frontier relative to τR |
Reading fidelity
high
Study strength
high
|
n=2
|
| Under a regularity condition on tail payoffs, the expected frontier overshoot beyond τR is at most O(1/ω), where ω is the monitoring/detection rate. Ai Safety And Ethics | negative | Expected excess technology frontier beyond the representative stopping threshold |
Reading fidelity
high
Study strength
high
|
n=2
O(1/ω) in expectation
|
| In the stationary environment with τR = 0, low trust in the rival's rationality implies that every equilibrium races to ruin and disaster occurs with probability one. Ai Safety And Ethics | negative | Probability of perpetual racing and disaster |
Reading fidelity
high
Study strength
high
|
n=2
|
| At intermediate levels of trust, both immediate stopping and racing to ruin are equilibria. Ai Safety And Ethics | mixed | Equilibrium selection between immediate stopping and perpetual racing |
Reading fidelity
high
Study strength
high
|
n=2
|
| At high trust, the probability that two rational firms race forever declines quadratically in the prior odds ratio of rationality. Ai Safety And Ethics | positive | Probability of perpetual racing by two rational firms |
Reading fidelity
high
Study strength
high
|
n=2
vanishes quadratically in the prior odds ratio of rationality
|
| Transparency is double-edged: at intermediate trust, increasing transparency can first eliminate the early-stopping equilibrium and later restore it when detection becomes sufficiently fast. Governance And Regulation | mixed | Existence of early-stopping equilibria as monitoring speed increases |
Reading fidelity
high
Study strength
high
|
n=2
|
| Increasing the diagonal marginal payoff from technology raises the representative stopping threshold τR, while increasing the disaster hazard lowers τR. Ai Safety And Ethics | mixed | Representative firm's technology stopping threshold |
Reading fidelity
high
Study strength
high
|
n=2
|