0 cumulative citations
View corpus contextFrontier models fail to cooperate in many high-stakes multi-agent tests: across 1,535 game-theoretic scenarios agents chose harmful actions 38% of the time. Simple game-theoretic prompt fixes raised cooperative outcomes by as much as 18%, indicating mitigations help but substantial reliability gaps remain.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
7 cumulative citations
View corpus contextFrontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.
Summary
Main Finding
GT-HARMBENCH is a large benchmark (1,535 high‑stakes scenarios) that maps real-world AI safety risks onto canonical 2×2 games. Across 15 frontier models, agents select the utilitarian/socially optimal action only 62% of the time (i.e., they fail 38% of the time). Failures arise from both conflict (e.g., Prisoner’s Dilemma) and coordination (e.g., Stag Hunt, Battle of the Sexes). Simple game‑theoretic interventions (mediation, communication, commitment devices) can improve socially beneficial outcomes by roughly 14–18%.
Key Points
- Scope and novelty
- First large-scale benchmark evaluating multi‑agent LLM safety in realistic, high‑stakes scenarios grounded in the MIT AI Risk Repository.
- 1,535 validated scenarios mapped into six canonical symmetric games: Prisoner’s Dilemma (490), Chicken (379), Stag Hunt (317), Coordination (180), Battle of the Sexes (141), No Conflict (28).
- Aggregate outcomes
- Utilitarian (total utility) accuracy across models: 62% (i.e., 38% socially suboptimal choices).
- Game‑level averages (utilitarian accuracy): Prisoner’s Dilemma 46%, Chicken 78%, Battle of the Sexes 46%, Stag Hunt 60%, Coordination 80%, No Conflict 100%.
- Example: mutual cooperation in Prisoner’s Dilemma only occurs 44% of the time.
- Model variation
- 15 frontier models evaluated (closed and open families). Aggregate ordering observed (Anthropic models top on average), but no monotonic relationship between standard capability measures and multi‑agent social welfare.
- Sensitivity to framing and ordering
- Presenting explicit numerical payoffs (making the game structure explicit) increases Nash‑equilibrium play (+6.2% Nash accuracy) but decreases utilitarian (social) accuracy (−4.06%), i.e., surfacing payoffs nudges models to more self‑interested equilibrium play.
- Randomizing option order in coordination tasks causes ~15% drop in coordination success — models rely on positional/focal cues.
- Interventions
- Mechanism design-style interventions (communication channels, commitments, mediation) yield improvements of ~14–18% in socially desirable outcomes; mediation performed best in experiments reported.
- Dataset quality checks
- Mapping: from 1,612 MIT risk entries, 604 (37.5%) identified as multi‑actor, producing 1,816 candidate (risk,game) pairs.
- Scenario generation and filtering used GPT‑5.1; 1,535 scenarios passed quality and game‑structure thresholds (84.5% pass rate).
- Human validation (30 sampled scenarios): inter‑annotator κ = 0.84 (86.7% raw agreement).
- Mechanical verification: 1,530/1,535 (99.7%) satisfy canonical ordinal conditions for their target game.
Data & Methods
- Data sources and mapping
- Base: MIT AI Risk Repository (Slattery et al., 2024).
- Mapping pipeline: GPT‑5.1 determines plausible canonical game(s) per risk (inclusive mapping: mean 3.01 games per risk).
- Scenario generation
- For each (risk, game) pair GPT‑5.1 produced first‑person contextualized scenarios with action labels, explicit numeric payoffs in [−10, 10], and a severity score (1–10).
- Filtering stage applied rubric scores (quality of contextualization and game‑correctness; retained if ≥8/10 on each).
- Structural choices
- Focus on symmetric 2×2 games (six canonical games) to isolate strategic tensions (cooperation vs. conflict vs. coordination) while keeping solution concepts well defined.
- Welfare metrics: utilitarian (primary reported), Rawlsian, and Nash social welfare — the analyses mainly use utilitarian welfare; occasional divergence (notably in Chicken).
- Evaluation
- Zero‑shot self‑play setup (each model plays both roles using its own policy) to avoid combinatorial cross‑play complexity; cross‑play results noted to typically increase miscoordination but relegated to appendix.
- 15 frontier models across major families (OpenAI GPT series, Anthropic Claude variants, Google Gemini, Grok, LLaMA3, Qwen3, DeepSeek).
- Reported outcome: fraction of scenarios where model choices jointly maximize the chosen welfare function.
- Limitations noted by authors
- Self‑play likely underestimates miscoordination in heterogeneous model populations (mixed-model settings).
- Use of symmetric 2×2 games abstracts away role asymmetries present in many economic contexts.
- Scenario generation and mapping relied on LLMs (GPT‑5.1), which can introduce biases in framing or payoff instantiation.
Implications for AI Economics
- Multi‑agent externalities matter: LLMs deployed in interacting economic systems can produce coordination failures and conflict externalities (arms races, market manipulations, systemic financial destabilization). Benchmarks that evaluate models in isolation miss these systemic risks.
- Incentive and mechanism design are central: simple mechanism interventions (mediation, commitment devices, communication protocols) can notably improve aggregate outcomes (14–18% gains). This suggests that designing institutional rules and incentives around AI agents can be an effective lever to mitigate harms in economic systems.
- Evaluation and procurement: standard capability metrics do not predict socially beneficial multi‑agent behavior. Economic actors (firms, regulators) should include multi‑agent safety tests (benchmarks like GT‑HARMBENCH) in model evaluation, auditing, and procurement decisions.
- Framing effects and transparency tradeoffs: surfacing explicit payoffs or strategic structure can push models toward Nash (self‑interested) behavior and away from utilitarian outcomes. In economic systems, revealing incentives or payoff structures could change agent behavior in undesired ways — policy design should consider whether transparency helps or hurts social objectives.
- Need for institution-building and governance: because LLMs can systematically prefer individually rational but collectively harmful actions, economic governance (regulation, verified commitments, third‑party mediators) is important to prevent arms‑race dynamics and coordination failures at scale.
- Research directions for AI economics
- Extend evaluation to asymmetric games and richer economic settings (auctions, markets, bargaining with power asymmetries).
- Study heterogeneous agent populations (mixed‑model cross‑play), market dynamics, repeated interaction, and learning dynamics to model longer‑run economic equilibria and instability risks.
- Formal analysis of interventions: cost–benefit and incentive compatibility of mediators, enforcement mechanisms, and market institutions when agents are LLMs.
- Integrate benchmark results into macro‑economic and financial stress tests that account for algorithmic agent interactions.
If you want, I can: - extract representative scenario examples mapped to specific economic harms (markets, financial institutions, auctions), or - produce a short checklist for policymakers and economists to operationalize these findings in procurement, regulation, or market design.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. Adoption Rate | positive | capability and deployment of frontier AI in high-stakes multi-agent settings |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. Ai Safety And Ethics | negative | coverage of multi-agent risks in AI safety benchmarks |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Ai Safety And Ethics | positive | number and composition of benchmark scenarios |
Reading fidelity
high
Study strength
high
|
n=1535
|
| Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Ai Safety And Ethics | positive | source/provenance of scenarios |
Reading fidelity
high
Study strength
high
|
n=1535
|
| Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. Decision Quality | negative | rate of failure to choose socially beneficial actions |
Reading fidelity
high
Study strength
medium
|
n=1535
38% of high-stakes cases
|
| We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. Decision Quality | mixed | sensitivity of model behavior to prompt framing and ordering; patterns of reasoning associated with failures |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Game-theoretic interventions improve socially beneficial outcomes by up to 18%. Decision Quality | positive | increase in socially beneficial outcomes due to interventions |
Reading fidelity
high
Study strength
medium
|
n=1535
up to 18%
|
| Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. Ai Safety And Ethics | mixed | existence of reliability gaps in multi-agent alignment and availability of a standardized testbed |
Reading fidelity
high
Study strength
medium
|
n=1535
|
| The benchmark and code are available at https://github.com/causalNLP/gt-harmbench. Ai Safety And Ethics | positive | public availability of benchmark and code |
Reading fidelity
high
Study strength
high
|
not reported
|