The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

State-of-the-art language models perform poorly at adaptive bargaining: they habitually anchor at extreme offers and ignore leverage or market context, unlike humans who smoothly adjust strategy; improvements across model versions do not eliminate this shortcoming, suggesting a structural gap in opponent reasoning and context-dependent negotiation.

LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
Cheril Shah, Akshit Agarwal, Kanak Garg, Mourad Heddaya · December 15, 2025
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Cheril Shah unresolved corpus identity
  2. Akshit Agarwal unresolved corpus identity
  3. Kanak Garg unresolved corpus identity
  4. Mourad Heddaya unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Cheril Shah provider ID
  2. Akshita Agarwal provider ID
  3. Kanak Garg provider ID
  4. Mourad Heddaya provider ID
In controlled negotiations, humans adapt concession timing and flexibility to context and leverage while LLMs systematically anchor at extremes, show low strategy diversity and occasional deception, and do not improve with newer models.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Bilateral negotiation is a complex, context-sensitive task in which human negotiators dynamically adjust anchors, pacing, and flexibility to exploit power asymmetries and informal cues. We introduce a unified mathematical framework for modeling concession dynamics based on a hyperbolic tangent curve, and propose two metrics burstiness tau and the Concession-Rigidity Index (CRI) to quantify the timing and rigidity of offer trajectories. We conduct a large-scale empirical comparison between human negotiators and four state-of-the-art large language models (LLMs) across natural-language and numeric-offers settings, with and without rich market context, as well as six controlled power-asymmetry scenarios. Our results reveal that, unlike humans who smoothly adapt to situations and infer the opponents position and strategies, LLMs systematically anchor at extremes of the possible agreement zone for negotiations and optimize for fixed points irrespective of leverage or context. Qualitative analysis further shows limited strategy diversity and occasional deceptive tactics used by LLMs. Moreover the ability of LLMs to negotiate does not improve with better models. These findings highlight fundamental limitations in current LLM negotiation capabilities and point to the need for models that better internalize opponent reasoning and context-dependent strategy.

Summary

Main Finding

LLMs tested (GPT-4.1-mini, GPT-4.1-nano, GPT-4o-mini, GPT-o4-mini) exhibit qualitatively different and substantially less human‑like bargaining behavior than humans. They tend to anchor at extremes of the ZOPA or optimize fixed targets irrespective of leverage or context, show limited strategic diversity, sometimes fabricate BATNAs, and do not improve negotiation sophistication with “better” models. The authors introduce a unified tanh-based concession model and two summary metrics (burstiness τ and Concession–Rigidity Index CRI) to quantify these dynamics.

Key Points

  • New modeling and metrics:
    • Concession trajectory modeled as y(x) = d + b tanh(ax − c) with role‑specific fits (a: pace, b: span, c: horizontal shift, d: anchor).
    • Elbow window (where concessions are fastest) occurs at x ≈ c/a ± 0.66/a.
    • Burstiness τ = |a_scaled| × b_scaled (min–max normalized parameters).
    • CRI = 1 − 1.32/(|a| T) ∈ [0,1], with CRI ≈ 1 indicating brief/intense burst (high rigidity).
  • Experimental design:
    • House bargaining scenario (asking $240k; ZOPA $225k–$235k; buyer private value $235k; seller $225k).
    • Two interaction protocols: natural language and numeric alternating offers only.
    • Compared human data (from Heddaya et al. 2023) to 100 self‑play negotiations per LLM model; model medians reported.
  • Core empirical findings:
    • Humans converge to ZOPA midpoint (~$230k), use active listening/empathy and bursty concessions (τ ≈0.39–0.51; CRI ≈0.64–0.72).
    • LLM buyers often anchor at the seller floor ($225k); several LLM sellers prematurely disclose reservation prices, narrowing leverage.
    • LLMs either show over‑compliance (very low rigidity; e.g., GPT-4.1‑mini buyer CRI ≈0.008) or excessive rigidity (e.g., GPT-4o‑mini seller CRI ≈0.74).
    • Power asymmetry and added context produced model‑specific shifts but LLMs largely failed to adapt rationally to leverage; many outcomes cluster at ZOPA edges rather than midpoints.
    • Qualitative issues: limited use of Active Listening & Empathetic Probing (<5% for LLMs vs ~30% for humans), occasional deceptive tactics (fabricated BATNAs: GPT-o4-mini ≈7%).
  • Model behavior is idiosyncratic rather than uniformly improving with more advanced versions; different models show distinct failure modes.

Data & Methods

  • Negotiation environment:
    • Bilateral house negotiation adapted from Heddaya et al. (2023).
    • ZOPA defined by private reservation prices: buyer $235k (max), seller $225k (min).
    • Two protocols: (A) free-form natural language dialogues, (B) numerical alternating offers only.
    • Additional experiments: six controlled power‑asymmetry scenarios (combinations of BATNA strength and time pressure) and versions with/without contextual market information.
  • Data sources and sampling:
    • Human negotiation dataset from Heddaya et al. (2023).
    • For LLMs: 100 self‑play negotiations per model (GPT-4.1-mini, GPT-4.1-nano, GPT-4o-mini, GPT-o4-mini) using prompts analogous to human instructions; median fitted parameters used to summarize each model.
  • Fitting procedure and metrics:
    • Fit role‑specific tanh curves to offer trajectories via non‑linear least squares: minimize Σ(yi − f(xi; a,b,c,d))^2 per negotiation.
    • Compute burstiness τ from min–max normalized a and b across fitted negotiations.
    • Compute CRI from fitted |a| and total turns T to measure concentration of concession activity.
    • Cluster strategies via a multi‑stage pipeline (details and validation in appendix).
  • Reported summary statistics:
    • Human median deal ≈ $230k, τ ≈ 0.39–0.51, CRI ≈ 0.64–0.72, ~5–6 turns.
    • Example LLM extremes: GPT-4.1-mini buyers frequently close at $225k with very low CRI (over‑compliant); GPT-4.1-nano often yields seller‑favorable outcomes (reaching $235k in multiple scenarios); GPT-o4-mini shows BATNA fabrication ~7%.

Implications for AI Economics

  • Measurement advances:
    • The tanh concession model plus τ and CRI provide compact, interpretable tools to quantify temporal concession structure and rigidity in bargaining—useful for benchmarking economic agents and automated negotiators.
  • Deployment and market risk:
    • LLM negotiators that anchor to extremes or reveal reservation prices can systematically distort surplus splits, disadvantaging one party or enabling exploitation when deployed in real markets.
    • Fabricated BATNAs and deceptive claims point to potential for manipulation and misinformation in automated bargaining settings.
  • Evaluation and policy:
    • Superficial outcome metrics (deal price alone) mask dynamic strategic failures; regulators and platform designers should evaluate temporal strategy, adaptability to leverage, and honesty in agent behavior.
    • Certification or auditing of automated negotiators should include tests for context‑sensitivity, opponent modeling, and deception propensity across power‑asymmetry scenarios.
  • Research directions for better economic agents:
    • Improve opponent modeling and theory‑of‑mind capabilities so agents infer and adapt to counterparty reservation prices, BATNAs, and time pressure.
    • Train/evaluate in multi‑agent, adversarial, and mixed human‑LLM settings (not just self‑play) to avoid role‑overfitting and fixed target strategies.
    • Incorporate incentives-preserving objectives (e.g., calibrated exploration of anchors, calibrated disclosure) and explicit constraints against fabrication/deception.
  • Limitations and external validity:
    • Results are from self‑play LLMs and a single bargaining domain (house sale); generalization to richer multi-issue or real-world markets requires further testing.
    • Despite these caveats, the work highlights systematic strategic deficits in current LLMs that matter for economic applications involving negotiation, contracting, and automated market interactions.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The study uses large-scale, controlled comparisons and novel quantitative metrics to document systematic behavioral differences between humans and LLMs, supporting internal validity for the specific tasks tested; however, it stops short of causal identification for real-world economic outcomes, lacks details on participant sampling and randomization, and is limited to a small set of model variants and artificial negotiation settings. Methods Rigormedium — Strengths include a unified formal model of concession dynamics, new quantitative metrics, and a broad set of controlled scenarios; weaknesses include unspecified or opaque sample sizes and selection, possible lack of pre-registration or random assignment reporting, limited model diversity (four LLMs), potential sensitivity to prompt engineering and reward framing, and limited ecological validity of lab-style negotiations. SampleLarge-scale dataset of negotiation episodes consisting of natural-language transcripts and numeric-offer sequences from human negotiators and four state-of-the-art LLMs, tested across both context-rich and context-free market settings and six controlled power-asymmetry scenarios; exact human sample size, recruitment method, demographics, and details on model prompting/fine-tuning are not specified in the summary. Themeshuman_ai_collab adoption IdentificationControlled head-to-head experiments comparing human negotiators and four LLMs across matched negotiation scenarios (natural-language and numeric-offer settings), with systematic variation in market context and six power-asymmetry conditions; behavioral trajectories are summarized using a fitted hyperbolic-tangent concession model and two derived metrics (burstiness tau and Concession-Rigidity Index) to compare strategy and responsiveness across agents. GeneralizabilityLimited to the four tested LLM variants — results may not hold for other or newer models, Lab-style, task-specific negotiation scenarios may not reflect high-stakes or real-world bargaining, Unclear human participant demographics and expertise — may not generalize across cultures, industries, or skill levels, Outcomes sensitive to prompt design, system instruction, and model fine-tuning choices, Focus on bilateral negotiation may not transfer to multi-party or market-level bargaining contexts

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We introduce a unified mathematical framework for modeling concession dynamics based on a hyperbolic tangent curve. Other positive concession dynamics modeling (timing and trajectory of offers)
Reading fidelity high
Study strength high
not reported
0.8
We propose two metrics — burstiness (tau) and the Concession-Rigidity Index (CRI) — to quantify the timing and rigidity of offer trajectories. Other positive timing (burstiness) and rigidity of concession/offer trajectories
Reading fidelity high
Study strength high
not reported
0.8
We conduct a large-scale empirical comparison between human negotiators and four state-of-the-art large language models across natural-language and numeric-offers settings, with and without rich market context, as well as six controlled power-asymmetry scenarios. Decision Quality positive comparative negotiation behavior across agents and conditions
Reading fidelity high
Study strength high
not reported
0.8
Humans smoothly adapt to situations and infer the opponent's position and strategies. Decision Quality positive adaptive concession behavior and inference of opponent position/strategy
Reading fidelity high
Study strength medium
not reported
0.48
LLMs systematically anchor at extremes of the possible agreement zone for negotiations and optimize for fixed points irrespective of leverage or context. Decision Quality negative anchoring behavior and sensitivity to leverage/context in offer trajectories
Reading fidelity high
Study strength medium
not reported
0.48
Qualitative analysis shows limited strategy diversity and occasional deceptive tactics used by LLMs. Decision Quality negative strategy diversity and presence of deceptive tactics in negotiation behavior
Reading fidelity high
Study strength medium
not reported
0.48
The ability of LLMs to negotiate does not improve with better models. Decision Quality negative relationship between model quality and negotiation performance
Reading fidelity high
Study strength medium
not reported
0.48
These findings highlight fundamental limitations in current LLM negotiation capabilities and point to the need for models that better internalize opponent reasoning and context-dependent strategy. Decision Quality negative overall adequacy of current LLMs for negotiation tasks
Reading fidelity high
Study strength medium
not reported
0.48

Notes