0 cumulative citations
View corpus contextIn lab ‘AI races’ competitors take shortcuts not because they are inherently risk-seeking but because rivals do and because they fall behind; varying the maximum private setback risk had little effect, suggesting competitive dynamics and early momentum, not individual risk attitudes, drive unsafe development.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.
Summary
Main Finding
Unsafe development in an idealised two-player AI-race experiment is driven less by individual risk preferences or the exogenous maximum private risk level and more by dynamic, strategic factors: opponent behaviour, early-action momentum, and relative race position. Participants are more likely to choose Unsafe after observing an opponent choose Unsafe, falling behind increases Unsafe choices, and early (first-round) Unsafe choices predict later Unsafe behaviour. A reduced evolutionary model with four simple strategies reproduces these patterns, showing how conditional unsafe strategies can be favoured by race dynamics.
Key Points
-
Experimental setup (high level)
- Paired participants repeatedly choose between Safe (sS = 1 step) and Unsafe (sU = 1.5 steps) development.
- Unsafe yields higher immediate payoff and faster progress but increases a participant’s private accumulated setback risk proportional to their fraction of Unsafe actions, capped at treatment-specific pmax_r.
- Game horizon: at least 5 rounds, then each subsequent round ends with probability p = 0.2 (uncertain horizon).
- Three treatments: pmax_r ∈ {0.1, 0.6, 0.9}.
- Risk preferences elicited via Eckel & Grossman (2008) task.
-
Pre-registered hypotheses and outcomes
- Hypothesis 1 (higher pmax_r → more Safe / less Unsafe) was not supported: no significant difference between pmax_r = 0.6 and 0.9; Unsafe choices remained frequent even at pmax_r = 0.9.
- Hypothesis 2 (elicited risk preferences predict Unsafe choices, stronger when pmax_r high) was not supported: elicited risk attitudes did not significantly predict Unsafe actions.
-
Main empirical (exploratory) findings
- Opponent’s previous action is a robust predictor: observing an opponent choose Unsafe substantially increases the probability of choosing Unsafe next round.
- Race position matters: being behind increases Unsafe propensity (fear-of-falling-behind mechanism); being ahead reduces it.
- First-round Unsafe choice predicts persistent later Unsafe behaviour.
- The participant’s own previous action is not a robust predictor once opponent action and race state are controlled for.
- Treatment (pmax_r) differences are better explained by how treatments change payoffs and thus strategic dynamics than by a simple direct treatment effect after round 1.
-
Reduced-strategy evolutionary explanation
- Four strategies: Always Safe (AS), Always Unsafe (AU), Conditionally Safe (CS; starts Safe), Conditionally Antisocial Safe (CAS; starts Unsafe then conditions).
- Evolutionary model reproduces qualitative treatment-level patterns: AU favored at low pmax_r, CAS at intermediate pmax_r, CS at high pmax_r. With higher behavioural noise (weak selection, higher mutation) the stationary distribution becomes more diffuse, matching experimental heterogeneity.
- Interpretation: conditional, responsive strategies and early aggressive moves can be evolutionarily (strategically) favoured by race dynamics even when individuals are not intrinsically risk-seeking.
Data & Methods
-
Participants and data
- Analysis used decisions from round 2 onward: N = 2,888 observations across 338 participants (172 pairs); standard errors clustered at pair level.
- Demographics (sex, age, nationality) and elicited risk-preference measures were included as covariates; risk-preference covariates did not predict Unsafe choices.
-
Task specifics
- Per-round payoffs: Unsafe gives higher immediate payoff and larger step advance (1.5 vs 1).
- Private setback risk accumulates with fraction of Unsafe choices and is bounded by pmax_r (0.1, 0.6, 0.9 depending on treatment).
- Uncertain game length: minimum 5 rounds, then geometric stopping with p = 0.2.
-
Statistical analyses
- Pre-registered analyses: mixed-effects models (details in Supplement), testing treatment and risk-preference effects.
- Exploratory analyses: cluster-robust logistic regressions predicting Unsafe choice at round t using opponent previous action, own previous action, relative race-step difference (∆S), first-round action indicator, demographic and risk covariates.
- Key regression result: positive and significant coefficient for opponent’s previous Unsafe action across specifications; negative coefficient for ∆S (ahead → less Unsafe) in full specs; first-round Unsafe significant predictor of later Unsafe.
-
Evolutionary model
- Reduced strategy space of four strategies chosen to reflect dominant empirical regularities.
- Evolutionary dynamics studied across selection (β) and mutation (µ) parameters.
- Reference parameterization: β = 2, µ ≈ 0.02; best-fit (to capture experimental noisiness): weak selection β = 0.01, µ = 0.05.
- Model maps changes in pmax_r to payoff changes that alter evolutionary success of strategies.
Implications for AI Economics
-
Strategic dynamics matter
- Models and empirical work on AI races should incorporate dynamic, conditional strategies, path dependence, and the role of early moves. Static analyses or explanations based on individual risk preferences alone may miss central drivers of unsafe race behaviour.
- Fear of falling behind and responsiveness to competitors can sustain unsafe development even when private downside risk is high.
-
Policy design
- Reducing competitive pressure and winner-take-all incentives can be more effective than focusing solely on individuals’ risk attitudes. Measures include:
- Coordination mechanisms or industry-wide safety standards that reduce incentives to defect (e.g., safety accords, delayed disclosure protocols).
- Institutional changes that reduce first-mover advantages (e.g., prize structures that reward safe deployment, regulation that limits unilateral deployments, licensing/approval regimes).
- Information-design interventions to limit escalation through mimicry (careful transparency practices that avoid triggering immediate reciprocation of unsafe moves).
- Mechanisms that mitigate “falling behind” incentives (e.g., shared timelines, cooperative research consortia, joint testing/safety sandboxes).
- Reducing competitive pressure and winner-take-all incentives can be more effective than focusing solely on individuals’ risk attitudes. Measures include:
-
Modeling and evaluation
- Policy evaluations and theoretical models should simulate conditional strategies and evolutionary dynamics (including noise), and consider how early momentum or single strategic moves can have outsize long-run effects.
- Interventions that change payoffs or the strategic meaning of being ahead/behind (not just changing individual risk preferences) are likely to shift equilibrium behaviour.
-
Cautions and limitations
- The experiment is intentionally idealised: two-player setup, laboratory incentives, specific step/payoff numbers, and private setback risk. Real-world AI development has many more actors, more complex information structures, heterogeneous capabilities, and social/publicly shared risks.
- Results are exploratory beyond preregistered hypotheses; robustness to broader environments and scaled settings remains to be tested.
- Policy prescriptions should be tested in richer models and field settings before broad deployment.
Overall, the study highlights that unsafe AI development can emerge from strategic interaction and path-dependence in races. Effective mitigation should therefore target competitive structures and coordination institutions, not only individual-level risk attitudes.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| There was no significant difference in the frequency of Unsafe choices between the maximum-risk treatments of 60% and 90%. Ai Safety And Ethics | null_result | Frequency of Unsafe development choices |
Reading fidelity
high
Study strength
high
|
t = −0.0101
|
| Elicited individual risk preferences did not significantly predict Unsafe choices. Ai Safety And Ethics | null_result | Probability or occurrence of Unsafe development choices |
Reading fidelity
high
Study strength
high
|
n=338
|
| Participants were substantially more likely to choose Unsafe development after observing their opponent choose Unsafe development in the previous round. Ai Safety And Ethics | positive | Choice of Unsafe development in the current round |
Reading fidelity
high
Study strength
high
|
n=338
β̂ = 0.640, p = 0.001; β̂ = 0.607, p = 0.002
|
| Being further ahead in the race reduced the probability of choosing Unsafe, while falling behind increased the incentive to choose Unsafe. Ai Safety And Ethics | mixed | Probability of choosing Unsafe development conditional on race position |
Reading fidelity
high
Study strength
medium
|
n=338
β̂ = −0.296, p = 0.048
|
| Choosing Unsafe development in the first round predicted choosing Unsafe development in later rounds. Ai Safety And Ethics | positive | Later-round Unsafe development choices |
Reading fidelity
high
Study strength
medium
|
n=338
|
| Unsafe choices occurred frequently across all maximum-risk treatments, including when the maximum private risk was 90%. Ai Safety And Ethics | positive | Frequency of Unsafe development choices |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The participant's own previous action was not a robust predictor of later Unsafe play once the opponent's previous action and race state were included. Ai Safety And Ethics | null_result | Later-round Unsafe development choice |
Reading fidelity
high
Study strength
high
|
n=338
|
| The reduced evolutionary model qualitatively reproduced the treatment-level pattern observed in the experiment, including a relatively small difference between the two higher-risk treatments and a stronger contrast with the low-risk treatment. Ai Safety And Ethics | positive | Frequency of Unsafe development choices predicted and observed across risk treatments |
Reading fidelity
high
Study strength
low
|
not reported
|
| In the evolutionary model, Always Unsafe was favored at low private risk, Conditionally Antisocial Safe was dominant at intermediate private risk, and Conditionally Safe dominated at the highest private risk. Ai Safety And Ethics | mixed | Evolutionary success or dominance of behavioral strategies |
Reading fidelity
high
Study strength
speculative
|
not reported
|