1 cumulative citations
View corpus contextState-of-the-art language models perform poorly at adaptive bargaining: they habitually anchor at extreme offers and ignore leverage or market context, unlike humans who smoothly adjust strategy; improvements across model versions do not eliminate this shortcoming, suggesting a structural gap in opponent reasoning and context-dependent negotiation.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Bilateral negotiation is a complex, context-sensitive task in which human negotiators dynamically adjust anchors, pacing, and flexibility to exploit power asymmetries and informal cues. We introduce a unified mathematical framework for modeling concession dynamics based on a hyperbolic tangent curve, and propose two metrics burstiness tau and the Concession-Rigidity Index (CRI) to quantify the timing and rigidity of offer trajectories. We conduct a large-scale empirical comparison between human negotiators and four state-of-the-art large language models (LLMs) across natural-language and numeric-offers settings, with and without rich market context, as well as six controlled power-asymmetry scenarios. Our results reveal that, unlike humans who smoothly adapt to situations and infer the opponents position and strategies, LLMs systematically anchor at extremes of the possible agreement zone for negotiations and optimize for fixed points irrespective of leverage or context. Qualitative analysis further shows limited strategy diversity and occasional deceptive tactics used by LLMs. Moreover the ability of LLMs to negotiate does not improve with better models. These findings highlight fundamental limitations in current LLM negotiation capabilities and point to the need for models that better internalize opponent reasoning and context-dependent strategy.
Summary
Main Finding
LLMs tested (GPT-4.1-mini, GPT-4.1-nano, GPT-4o-mini, GPT-o4-mini) exhibit qualitatively different and substantially less human‑like bargaining behavior than humans. They tend to anchor at extremes of the ZOPA or optimize fixed targets irrespective of leverage or context, show limited strategic diversity, sometimes fabricate BATNAs, and do not improve negotiation sophistication with “better” models. The authors introduce a unified tanh-based concession model and two summary metrics (burstiness τ and Concession–Rigidity Index CRI) to quantify these dynamics.
Key Points
- New modeling and metrics:
- Concession trajectory modeled as y(x) = d + b tanh(ax − c) with role‑specific fits (a: pace, b: span, c: horizontal shift, d: anchor).
- Elbow window (where concessions are fastest) occurs at x ≈ c/a ± 0.66/a.
- Burstiness τ = |a_scaled| × b_scaled (min–max normalized parameters).
- CRI = 1 − 1.32/(|a| T) ∈ [0,1], with CRI ≈ 1 indicating brief/intense burst (high rigidity).
- Experimental design:
- House bargaining scenario (asking $240k; ZOPA $225k–$235k; buyer private value $235k; seller $225k).
- Two interaction protocols: natural language and numeric alternating offers only.
- Compared human data (from Heddaya et al. 2023) to 100 self‑play negotiations per LLM model; model medians reported.
- Core empirical findings:
- Humans converge to ZOPA midpoint (~$230k), use active listening/empathy and bursty concessions (τ ≈0.39–0.51; CRI ≈0.64–0.72).
- LLM buyers often anchor at the seller floor ($225k); several LLM sellers prematurely disclose reservation prices, narrowing leverage.
- LLMs either show over‑compliance (very low rigidity; e.g., GPT-4.1‑mini buyer CRI ≈0.008) or excessive rigidity (e.g., GPT-4o‑mini seller CRI ≈0.74).
- Power asymmetry and added context produced model‑specific shifts but LLMs largely failed to adapt rationally to leverage; many outcomes cluster at ZOPA edges rather than midpoints.
- Qualitative issues: limited use of Active Listening & Empathetic Probing (<5% for LLMs vs ~30% for humans), occasional deceptive tactics (fabricated BATNAs: GPT-o4-mini ≈7%).
- Model behavior is idiosyncratic rather than uniformly improving with more advanced versions; different models show distinct failure modes.
Data & Methods
- Negotiation environment:
- Bilateral house negotiation adapted from Heddaya et al. (2023).
- ZOPA defined by private reservation prices: buyer $235k (max), seller $225k (min).
- Two protocols: (A) free-form natural language dialogues, (B) numerical alternating offers only.
- Additional experiments: six controlled power‑asymmetry scenarios (combinations of BATNA strength and time pressure) and versions with/without contextual market information.
- Data sources and sampling:
- Human negotiation dataset from Heddaya et al. (2023).
- For LLMs: 100 self‑play negotiations per model (GPT-4.1-mini, GPT-4.1-nano, GPT-4o-mini, GPT-o4-mini) using prompts analogous to human instructions; median fitted parameters used to summarize each model.
- Fitting procedure and metrics:
- Fit role‑specific tanh curves to offer trajectories via non‑linear least squares: minimize Σ(yi − f(xi; a,b,c,d))^2 per negotiation.
- Compute burstiness τ from min–max normalized a and b across fitted negotiations.
- Compute CRI from fitted |a| and total turns T to measure concentration of concession activity.
- Cluster strategies via a multi‑stage pipeline (details and validation in appendix).
- Reported summary statistics:
- Human median deal ≈ $230k, τ ≈ 0.39–0.51, CRI ≈ 0.64–0.72, ~5–6 turns.
- Example LLM extremes: GPT-4.1-mini buyers frequently close at $225k with very low CRI (over‑compliant); GPT-4.1-nano often yields seller‑favorable outcomes (reaching $235k in multiple scenarios); GPT-o4-mini shows BATNA fabrication ~7%.
Implications for AI Economics
- Measurement advances:
- The tanh concession model plus τ and CRI provide compact, interpretable tools to quantify temporal concession structure and rigidity in bargaining—useful for benchmarking economic agents and automated negotiators.
- Deployment and market risk:
- LLM negotiators that anchor to extremes or reveal reservation prices can systematically distort surplus splits, disadvantaging one party or enabling exploitation when deployed in real markets.
- Fabricated BATNAs and deceptive claims point to potential for manipulation and misinformation in automated bargaining settings.
- Evaluation and policy:
- Superficial outcome metrics (deal price alone) mask dynamic strategic failures; regulators and platform designers should evaluate temporal strategy, adaptability to leverage, and honesty in agent behavior.
- Certification or auditing of automated negotiators should include tests for context‑sensitivity, opponent modeling, and deception propensity across power‑asymmetry scenarios.
- Research directions for better economic agents:
- Improve opponent modeling and theory‑of‑mind capabilities so agents infer and adapt to counterparty reservation prices, BATNAs, and time pressure.
- Train/evaluate in multi‑agent, adversarial, and mixed human‑LLM settings (not just self‑play) to avoid role‑overfitting and fixed target strategies.
- Incorporate incentives-preserving objectives (e.g., calibrated exploration of anchors, calibrated disclosure) and explicit constraints against fabrication/deception.
- Limitations and external validity:
- Results are from self‑play LLMs and a single bargaining domain (house sale); generalization to richer multi-issue or real-world markets requires further testing.
- Despite these caveats, the work highlights systematic strategic deficits in current LLMs that matter for economic applications involving negotiation, contracting, and automated market interactions.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We introduce a unified mathematical framework for modeling concession dynamics based on a hyperbolic tangent curve. Other | positive | concession dynamics modeling (timing and trajectory of offers) |
Reading fidelity
high
Study strength
high
|
not reported
|
| We propose two metrics — burstiness (tau) and the Concession-Rigidity Index (CRI) — to quantify the timing and rigidity of offer trajectories. Other | positive | timing (burstiness) and rigidity of concession/offer trajectories |
Reading fidelity
high
Study strength
high
|
not reported
|
| We conduct a large-scale empirical comparison between human negotiators and four state-of-the-art large language models across natural-language and numeric-offers settings, with and without rich market context, as well as six controlled power-asymmetry scenarios. Decision Quality | positive | comparative negotiation behavior across agents and conditions |
Reading fidelity
high
Study strength
high
|
not reported
|
| Humans smoothly adapt to situations and infer the opponent's position and strategies. Decision Quality | positive | adaptive concession behavior and inference of opponent position/strategy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLMs systematically anchor at extremes of the possible agreement zone for negotiations and optimize for fixed points irrespective of leverage or context. Decision Quality | negative | anchoring behavior and sensitivity to leverage/context in offer trajectories |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Qualitative analysis shows limited strategy diversity and occasional deceptive tactics used by LLMs. Decision Quality | negative | strategy diversity and presence of deceptive tactics in negotiation behavior |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The ability of LLMs to negotiate does not improve with better models. Decision Quality | negative | relationship between model quality and negotiation performance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| These findings highlight fundamental limitations in current LLM negotiation capabilities and point to the need for models that better internalize opponent reasoning and context-dependent strategy. Decision Quality | negative | overall adequacy of current LLMs for negotiation tasks |
Reading fidelity
high
Study strength
medium
|
not reported
|