The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Frontier models fail to cooperate in many high-stakes multi-agent tests: across 1,535 game-theoretic scenarios agents chose harmful actions 38% of the time. Simple game-theoretic prompt fixes raised cooperative outcomes by as much as 18%, indicating mitigations help but substantial reliability gaps remain.

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
Cobben, Pepijn, Huang, Xuanqiang Angelo, Pham, Thao Amelia, Dahlgren, Isabel, Zhang, Terry Jingchen, Jin, Zhijing · February 12, 2026 · arXiv (Cornell University)
openalex descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Cobben, Pepijn provider ID
  2. Huang, Xuanqiang Angelo provider ID
  3. Pham, Thao Amelia provider ID
  4. Dahlgren, Isabel provider ID
  5. Zhang, Terry Jingchen provider ID
  6. Jin, Zhijing provider ID

Semantic Scholar

Latest observation:

  1. Pepijn Cobben provider ID
  2. Xuan Huang provider ID
  3. Thao Pham provider ID
  4. Isabel Dahlgren provider ID
  5. Terry Zhang provider ID
  6. Zhijing Jin provider ID
In a benchmark of 1,535 high-stakes game-theoretic scenarios, 15 frontier models selected non-socially-beneficial actions in 38% of cases, while targeted game-theoretic prompt interventions improved cooperative outcomes by up to 18%.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.

Summary

Main Finding

GT-HARMBENCH is a large benchmark (1,535 high‑stakes scenarios) that maps real-world AI safety risks onto canonical 2×2 games. Across 15 frontier models, agents select the utilitarian/socially optimal action only 62% of the time (i.e., they fail 38% of the time). Failures arise from both conflict (e.g., Prisoner’s Dilemma) and coordination (e.g., Stag Hunt, Battle of the Sexes). Simple game‑theoretic interventions (mediation, communication, commitment devices) can improve socially beneficial outcomes by roughly 14–18%.

Key Points

  • Scope and novelty
    • First large-scale benchmark evaluating multi‑agent LLM safety in realistic, high‑stakes scenarios grounded in the MIT AI Risk Repository.
    • 1,535 validated scenarios mapped into six canonical symmetric games: Prisoner’s Dilemma (490), Chicken (379), Stag Hunt (317), Coordination (180), Battle of the Sexes (141), No Conflict (28).
  • Aggregate outcomes
    • Utilitarian (total utility) accuracy across models: 62% (i.e., 38% socially suboptimal choices).
    • Game‑level averages (utilitarian accuracy): Prisoner’s Dilemma 46%, Chicken 78%, Battle of the Sexes 46%, Stag Hunt 60%, Coordination 80%, No Conflict 100%.
    • Example: mutual cooperation in Prisoner’s Dilemma only occurs 44% of the time.
  • Model variation
    • 15 frontier models evaluated (closed and open families). Aggregate ordering observed (Anthropic models top on average), but no monotonic relationship between standard capability measures and multi‑agent social welfare.
  • Sensitivity to framing and ordering
    • Presenting explicit numerical payoffs (making the game structure explicit) increases Nash‑equilibrium play (+6.2% Nash accuracy) but decreases utilitarian (social) accuracy (−4.06%), i.e., surfacing payoffs nudges models to more self‑interested equilibrium play.
    • Randomizing option order in coordination tasks causes ~15% drop in coordination success — models rely on positional/focal cues.
  • Interventions
    • Mechanism design-style interventions (communication channels, commitments, mediation) yield improvements of ~14–18% in socially desirable outcomes; mediation performed best in experiments reported.
  • Dataset quality checks
    • Mapping: from 1,612 MIT risk entries, 604 (37.5%) identified as multi‑actor, producing 1,816 candidate (risk,game) pairs.
    • Scenario generation and filtering used GPT‑5.1; 1,535 scenarios passed quality and game‑structure thresholds (84.5% pass rate).
    • Human validation (30 sampled scenarios): inter‑annotator κ = 0.84 (86.7% raw agreement).
    • Mechanical verification: 1,530/1,535 (99.7%) satisfy canonical ordinal conditions for their target game.

Data & Methods

  • Data sources and mapping
    • Base: MIT AI Risk Repository (Slattery et al., 2024).
    • Mapping pipeline: GPT‑5.1 determines plausible canonical game(s) per risk (inclusive mapping: mean 3.01 games per risk).
  • Scenario generation
    • For each (risk, game) pair GPT‑5.1 produced first‑person contextualized scenarios with action labels, explicit numeric payoffs in [−10, 10], and a severity score (1–10).
    • Filtering stage applied rubric scores (quality of contextualization and game‑correctness; retained if ≥8/10 on each).
  • Structural choices
    • Focus on symmetric 2×2 games (six canonical games) to isolate strategic tensions (cooperation vs. conflict vs. coordination) while keeping solution concepts well defined.
    • Welfare metrics: utilitarian (primary reported), Rawlsian, and Nash social welfare — the analyses mainly use utilitarian welfare; occasional divergence (notably in Chicken).
  • Evaluation
    • Zero‑shot self‑play setup (each model plays both roles using its own policy) to avoid combinatorial cross‑play complexity; cross‑play results noted to typically increase miscoordination but relegated to appendix.
    • 15 frontier models across major families (OpenAI GPT series, Anthropic Claude variants, Google Gemini, Grok, LLaMA3, Qwen3, DeepSeek).
    • Reported outcome: fraction of scenarios where model choices jointly maximize the chosen welfare function.
  • Limitations noted by authors
    • Self‑play likely underestimates miscoordination in heterogeneous model populations (mixed-model settings).
    • Use of symmetric 2×2 games abstracts away role asymmetries present in many economic contexts.
    • Scenario generation and mapping relied on LLMs (GPT‑5.1), which can introduce biases in framing or payoff instantiation.

Implications for AI Economics

  • Multi‑agent externalities matter: LLMs deployed in interacting economic systems can produce coordination failures and conflict externalities (arms races, market manipulations, systemic financial destabilization). Benchmarks that evaluate models in isolation miss these systemic risks.
  • Incentive and mechanism design are central: simple mechanism interventions (mediation, commitment devices, communication protocols) can notably improve aggregate outcomes (14–18% gains). This suggests that designing institutional rules and incentives around AI agents can be an effective lever to mitigate harms in economic systems.
  • Evaluation and procurement: standard capability metrics do not predict socially beneficial multi‑agent behavior. Economic actors (firms, regulators) should include multi‑agent safety tests (benchmarks like GT‑HARMBENCH) in model evaluation, auditing, and procurement decisions.
  • Framing effects and transparency tradeoffs: surfacing explicit payoffs or strategic structure can push models toward Nash (self‑interested) behavior and away from utilitarian outcomes. In economic systems, revealing incentives or payoff structures could change agent behavior in undesired ways — policy design should consider whether transparency helps or hurts social objectives.
  • Need for institution-building and governance: because LLMs can systematically prefer individually rational but collectively harmful actions, economic governance (regulation, verified commitments, third‑party mediators) is important to prevent arms‑race dynamics and coordination failures at scale.
  • Research directions for AI economics
    • Extend evaluation to asymmetric games and richer economic settings (auctions, markets, bargaining with power asymmetries).
    • Study heterogeneous agent populations (mixed‑model cross‑play), market dynamics, repeated interaction, and learning dynamics to model longer‑run economic equilibria and instability risks.
    • Formal analysis of interventions: cost–benefit and incentive compatibility of mediators, enforcement mechanisms, and market institutions when agents are LLMs.
    • Integrate benchmark results into macro‑economic and financial stress tests that account for algorithmic agent interactions.

If you want, I can: - extract representative scenario examples mapped to specific economic harms (markets, financial institutions, auctions), or - produce a short checklist for policymakers and economists to operationalize these findings in procurement, regulation, or market design.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The dataset is large and covers many standardized game-theoretic scenarios and multiple models, so findings about model tendencies are well supported within the benchmark; however, the scenarios are synthetic/constructed, outcome coding (what counts as 'socially beneficial') may involve researcher judgment, and results do not directly translate to measured economic outcomes in real-world multi-agent systems. Methods Rigormedium — The authors systematically vary prompts, ordering, and interventions across many scenarios and models and report aggregate failure rates and improvement magnitudes, but potential concerns include scenario selection bias, subjectivity in labeling outcomes, lack of dynamic multi-turn institutional context, and limited transparency on model selection and statistical uncertainty for some comparisons. Sample1,535 high-stakes, game-theoretic scenarios (Prisoner's Dilemma, Stag Hunt, Chicken, etc.) drawn from the MIT AI Risk Repository covering contexts like military escalation, election manipulation, and medical malpractice; evaluated across 15 frontier language/decision models with multiple prompt framings, orderings, and a set of game-theoretic interventions; outcomes coded for whether agents chose 'socially beneficial' actions. Themesgovernance human_ai_collab adoption IdentificationBehavioral benchmarking using a large set of constructed, game-theoretic scenarios: compare model choices across 1,535 controlled scenarios drawn from the MIT AI Risk Repository and measure changes under alternative prompt framings and explicit game-theoretic interventions; no quasi-experimental or instrumental-variable identification for real-world causal effects. GeneralizabilityScenarios are constructed and may not capture full complexity of real-world institutions, repeated interactions, or incentives, Findings limited to the specific 15 frontier models tested and prompt formats used, Outcome definitions (socially beneficial vs harmful) may be subjective and context-dependent, Static, single- or few-shot prompts may not reflect long-run multi-agent dynamics or deployment environments, Benchmarks emphasize game-theoretic structures and may not generalize to other multi-agent settings (market interactions, hierarchical organizations), Limited information on model access modes (API vs fine-tuned) and temperature/decoding settings could affect replicability

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. Adoption Rate positive capability and deployment of frontier AI in high-stakes multi-agent settings
Reading fidelity high
Study strength medium
not reported
0.18
Existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. Ai Safety And Ethics negative coverage of multi-agent risks in AI safety benchmarks
Reading fidelity high
Study strength medium
not reported
0.18
We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Ai Safety And Ethics positive number and composition of benchmark scenarios
Reading fidelity high
Study strength high
n=1535
0.3
Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Ai Safety And Ethics positive source/provenance of scenarios
Reading fidelity high
Study strength high
n=1535
0.3
Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. Decision Quality negative rate of failure to choose socially beneficial actions
Reading fidelity high
Study strength medium
n=1535
38% of high-stakes cases
0.18
We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. Decision Quality mixed sensitivity of model behavior to prompt framing and ordering; patterns of reasoning associated with failures
Reading fidelity high
Study strength medium
not reported
0.18
Game-theoretic interventions improve socially beneficial outcomes by up to 18%. Decision Quality positive increase in socially beneficial outcomes due to interventions
Reading fidelity high
Study strength medium
n=1535
up to 18%
0.18
Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. Ai Safety And Ethics mixed existence of reliability gaps in multi-agent alignment and availability of a standardized testbed
Reading fidelity high
Study strength medium
n=1535
0.18
The benchmark and code are available at https://github.com/causalNLP/gt-harmbench. Ai Safety And Ethics positive public availability of benchmark and code
Reading fidelity high
Study strength high
not reported
0.3

Notes