The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

In lab ‘AI races’ competitors take shortcuts not because they are inherently risk-seeking but because rivals do and because they fall behind; varying the maximum private setback risk had little effect, suggesting competitive dynamics and early momentum, not individual risk attitudes, drive unsafe development.

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
Elias Fernández Domingos, The Anh Han · July 28, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Elias Fernández Domingos unresolved corpus identity
  2. The Anh Han unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Elias Fernández Domingos provider ID
  2. T. Han provider ID
In a preregistered framed experiment, participants’ unsafe choices in an idealised two-player AI race were driven more by opponent behaviour, early-round momentum, and fear of falling behind than by the exogenously manipulated maximum private risk or individual risk preferences.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.

Summary

Main Finding

Unsafe development in an idealised two-player AI-race experiment is driven less by individual risk preferences or the exogenous maximum private risk level and more by dynamic, strategic factors: opponent behaviour, early-action momentum, and relative race position. Participants are more likely to choose Unsafe after observing an opponent choose Unsafe, falling behind increases Unsafe choices, and early (first-round) Unsafe choices predict later Unsafe behaviour. A reduced evolutionary model with four simple strategies reproduces these patterns, showing how conditional unsafe strategies can be favoured by race dynamics.

Key Points

  • Experimental setup (high level)

    • Paired participants repeatedly choose between Safe (sS = 1 step) and Unsafe (sU = 1.5 steps) development.
    • Unsafe yields higher immediate payoff and faster progress but increases a participant’s private accumulated setback risk proportional to their fraction of Unsafe actions, capped at treatment-specific pmax_r.
    • Game horizon: at least 5 rounds, then each subsequent round ends with probability p = 0.2 (uncertain horizon).
    • Three treatments: pmax_r ∈ {0.1, 0.6, 0.9}.
    • Risk preferences elicited via Eckel & Grossman (2008) task.
  • Pre-registered hypotheses and outcomes

    • Hypothesis 1 (higher pmax_r → more Safe / less Unsafe) was not supported: no significant difference between pmax_r = 0.6 and 0.9; Unsafe choices remained frequent even at pmax_r = 0.9.
    • Hypothesis 2 (elicited risk preferences predict Unsafe choices, stronger when pmax_r high) was not supported: elicited risk attitudes did not significantly predict Unsafe actions.
  • Main empirical (exploratory) findings

    • Opponent’s previous action is a robust predictor: observing an opponent choose Unsafe substantially increases the probability of choosing Unsafe next round.
    • Race position matters: being behind increases Unsafe propensity (fear-of-falling-behind mechanism); being ahead reduces it.
    • First-round Unsafe choice predicts persistent later Unsafe behaviour.
    • The participant’s own previous action is not a robust predictor once opponent action and race state are controlled for.
    • Treatment (pmax_r) differences are better explained by how treatments change payoffs and thus strategic dynamics than by a simple direct treatment effect after round 1.
  • Reduced-strategy evolutionary explanation

    • Four strategies: Always Safe (AS), Always Unsafe (AU), Conditionally Safe (CS; starts Safe), Conditionally Antisocial Safe (CAS; starts Unsafe then conditions).
    • Evolutionary model reproduces qualitative treatment-level patterns: AU favored at low pmax_r, CAS at intermediate pmax_r, CS at high pmax_r. With higher behavioural noise (weak selection, higher mutation) the stationary distribution becomes more diffuse, matching experimental heterogeneity.
    • Interpretation: conditional, responsive strategies and early aggressive moves can be evolutionarily (strategically) favoured by race dynamics even when individuals are not intrinsically risk-seeking.

Data & Methods

  • Participants and data

    • Analysis used decisions from round 2 onward: N = 2,888 observations across 338 participants (172 pairs); standard errors clustered at pair level.
    • Demographics (sex, age, nationality) and elicited risk-preference measures were included as covariates; risk-preference covariates did not predict Unsafe choices.
  • Task specifics

    • Per-round payoffs: Unsafe gives higher immediate payoff and larger step advance (1.5 vs 1).
    • Private setback risk accumulates with fraction of Unsafe choices and is bounded by pmax_r (0.1, 0.6, 0.9 depending on treatment).
    • Uncertain game length: minimum 5 rounds, then geometric stopping with p = 0.2.
  • Statistical analyses

    • Pre-registered analyses: mixed-effects models (details in Supplement), testing treatment and risk-preference effects.
    • Exploratory analyses: cluster-robust logistic regressions predicting Unsafe choice at round t using opponent previous action, own previous action, relative race-step difference (∆S), first-round action indicator, demographic and risk covariates.
    • Key regression result: positive and significant coefficient for opponent’s previous Unsafe action across specifications; negative coefficient for ∆S (ahead → less Unsafe) in full specs; first-round Unsafe significant predictor of later Unsafe.
  • Evolutionary model

    • Reduced strategy space of four strategies chosen to reflect dominant empirical regularities.
    • Evolutionary dynamics studied across selection (β) and mutation (µ) parameters.
    • Reference parameterization: β = 2, µ ≈ 0.02; best-fit (to capture experimental noisiness): weak selection β = 0.01, µ = 0.05.
    • Model maps changes in pmax_r to payoff changes that alter evolutionary success of strategies.

Implications for AI Economics

  • Strategic dynamics matter

    • Models and empirical work on AI races should incorporate dynamic, conditional strategies, path dependence, and the role of early moves. Static analyses or explanations based on individual risk preferences alone may miss central drivers of unsafe race behaviour.
    • Fear of falling behind and responsiveness to competitors can sustain unsafe development even when private downside risk is high.
  • Policy design

    • Reducing competitive pressure and winner-take-all incentives can be more effective than focusing solely on individuals’ risk attitudes. Measures include:
      • Coordination mechanisms or industry-wide safety standards that reduce incentives to defect (e.g., safety accords, delayed disclosure protocols).
      • Institutional changes that reduce first-mover advantages (e.g., prize structures that reward safe deployment, regulation that limits unilateral deployments, licensing/approval regimes).
      • Information-design interventions to limit escalation through mimicry (careful transparency practices that avoid triggering immediate reciprocation of unsafe moves).
      • Mechanisms that mitigate “falling behind” incentives (e.g., shared timelines, cooperative research consortia, joint testing/safety sandboxes).
  • Modeling and evaluation

    • Policy evaluations and theoretical models should simulate conditional strategies and evolutionary dynamics (including noise), and consider how early momentum or single strategic moves can have outsize long-run effects.
    • Interventions that change payoffs or the strategic meaning of being ahead/behind (not just changing individual risk preferences) are likely to shift equilibrium behaviour.
  • Cautions and limitations

    • The experiment is intentionally idealised: two-player setup, laboratory incentives, specific step/payoff numbers, and private setback risk. Real-world AI development has many more actors, more complex information structures, heterogeneous capabilities, and social/publicly shared risks.
    • Results are exploratory beyond preregistered hypotheses; robustness to broader environments and scaled settings remains to be tested.
    • Policy prescriptions should be tested in richer models and field settings before broad deployment.

Overall, the study highlights that unsafe AI development can emerge from strategic interaction and path-dependence in races. Effective mitigation should therefore target competitive structures and coordination institutions, not only individual-level risk attitudes.

Assessment

Paper Typerct Evidence Strengthmedium — The randomized experimental design gives good internal validity for causal comparisons of the exogenous treatment (p_max_r). However, the main substantive conclusions rely largely on exploratory, within-interaction dynamics (responses to opponent actions, first-round momentum, and relative position), which are observational conditional relations within the repeated game rather than direct randomized manipulations. Sample and ecological limitations (lab/online framed task, small-scale two-player setting, artificially low stakes relative to real-world AI actors) constrain external validity. Methods Rigormedium — Strengths include pre-registration, randomized treatment assignment, appropriate clustering of standard errors at the pair level, use of risk-elicitations as covariates, and a theory-driven evolutionary model to interpret patterns. Weaknesses: key claims are supported mainly by exploratory analyses rather than preregistered tests; potential multiple-hypothesis concerns and post-hoc model choices; limited discussion (in supplied text) of power/sample-size justification and of robustness checks for dynamic dependence and learning; external validity issues inherent to framed experiments. SampleParticipants matched in pairs in a framed two-player repeated 'AI race' experiment; reported analysis uses N = 338 participants (172 pairs) and 2,888 round-level observations from round 2 onward; between-subjects manipulation of maximum private setback risk p_max_r ∈ {0.1, 0.6, 0.9}; risk preferences elicited using an Eckel and Grossman (2008) task; demographics recorded (sex, age, nationality). (Setting — lab vs online — not explicit in the provided excerpt.) Themesgovernance innovation IdentificationBetween-subjects randomized treatment manipulating the maximum private risk (p_max_r ∈ {0.1, 0.6, 0.9}) in a framed two-player repeated behavioural experiment; pre-registered primary hypotheses; analysis uses cluster-robust logistic regressions and mixed-effects models to compare unsafe-choice frequencies across treatments and to relate choices to elicited risk preferences and dynamic game-state variables (previous actions, relative position). A complementary evolutionary game-theory model is used to interpret observed behavioural patterns. GeneralizabilityFramed, idealised two-player lab/online task with low monetary/stakes compared with firms or states limits external validity to real-world AI developers., Behavior of individual participants in short repeated games may not map to organizational decision-making, institutional constraints, or multi-party international races., Sample likely non-representative (convenience sample typical of behavioural experiments); cultural and incentive differences may matter., Simplified payoff and risk accumulation rules abstract from complex, multi-dimensional safety risks, regulation, and information asymmetries in real AI development.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
There was no significant difference in the frequency of Unsafe choices between the maximum-risk treatments of 60% and 90%. Ai Safety And Ethics null_result Frequency of Unsafe development choices
Reading fidelity high
Study strength high
t = −0.0101
1.0
Elicited individual risk preferences did not significantly predict Unsafe choices. Ai Safety And Ethics null_result Probability or occurrence of Unsafe development choices
Reading fidelity high
Study strength high
n=338
1.0
Participants were substantially more likely to choose Unsafe development after observing their opponent choose Unsafe development in the previous round. Ai Safety And Ethics positive Choice of Unsafe development in the current round
Reading fidelity high
Study strength high
n=338
β̂ = 0.640, p = 0.001; β̂ = 0.607, p = 0.002
1.0
Being further ahead in the race reduced the probability of choosing Unsafe, while falling behind increased the incentive to choose Unsafe. Ai Safety And Ethics mixed Probability of choosing Unsafe development conditional on race position
Reading fidelity high
Study strength medium
n=338
β̂ = −0.296, p = 0.048
0.6
Choosing Unsafe development in the first round predicted choosing Unsafe development in later rounds. Ai Safety And Ethics positive Later-round Unsafe development choices
Reading fidelity high
Study strength medium
n=338
0.6
Unsafe choices occurred frequently across all maximum-risk treatments, including when the maximum private risk was 90%. Ai Safety And Ethics positive Frequency of Unsafe development choices
Reading fidelity high
Study strength medium
not reported
0.6
The participant's own previous action was not a robust predictor of later Unsafe play once the opponent's previous action and race state were included. Ai Safety And Ethics null_result Later-round Unsafe development choice
Reading fidelity high
Study strength high
n=338
1.0
The reduced evolutionary model qualitatively reproduced the treatment-level pattern observed in the experiment, including a relatively small difference between the two higher-risk treatments and a stronger contrast with the low-risk treatment. Ai Safety And Ethics positive Frequency of Unsafe development choices predicted and observed across risk treatments
Reading fidelity high
Study strength low
not reported
0.3
In the evolutionary model, Always Unsafe was favored at low private risk, Conditionally Antisocial Safe was dominant at intermediate private risk, and Conditionally Safe dominated at the highest private risk. Ai Safety And Ethics mixed Evolutionary success or dominance of behavioral strategies
Reading fidelity high
Study strength speculative
not reported
0.1

Notes