1 cumulative citations
View corpus contextLanguage alters LLM negotiation outcomes more than model changes: switching from English to several Indic languages can reverse proposer advantages and shift surplus allocation; distributive games become less stable while integrative games induce more exploratory deals.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Negotiation is a core component of social intelligence, requiring agents to balance strategic reasoning, cooperation, and social norms. Recent work shows that LLMs can engage in multi-turn negotiation, yet nearly all evaluations occur exclusively in English. Using controlled multi-agent simulations across Ultimatum, Buy-Sell, and Resource Exchange games, we systematically isolate language effects across English and four Indic framings (Hindi, Punjabi, Gujarati, Marwadi) by holding game rules, model parameters, and incentives constant across all conditions. We find that language choice can shift outcomes more strongly than changing models, reversing proposer advantages and reallocating surplus. Crucially, effects are task-contingent: Indic languages reduce stability in distributive games yet induce richer exploration in integrative settings. Our results demonstrate that evaluating LLM negotiation solely in English yields incomplete and potentially misleading conclusions. These findings caution against English-only evaluation of LLMs and suggest that culturally-aware evaluation is essential for fair deployment.
Summary
Main Finding
The interaction language systematically alters LLM negotiation behavior. Holding game rules, incentives, and model parameters constant across 4,320 simulated games, the authors show language choice acts as a strategic prior that changes outcomes and dynamics in task-dependent ways: Indic framings (Hindi, Gujarati, Punjabi) reduce stability in distributive games (Ultimatum, Buy–Sell) — longer negotiations and lower acceptance in some cases — while increasing exploratory trade (higher trade volumes) in an integrative Resource Exchange game. These effects persist across multiple LLM architectures and are robust to prompt variants and script choice.
Key Points
- Experiments: 4,320 multi-agent games covering three canonical negotiation settings (Ultimatum, Buy–Sell, Resource Exchange), four language conditions (English baseline; Hindi, Gujarati, Punjabi), and four multilingual LLMs (GPT-4o, GPT-3.5 Turbo, Claude-3-Haiku, Claude-3.5-Haiku).
- Statistically significant language effects (nonparametric tests; Benjamini–Hochberg correction):
- Ultimatum: English acceptance rate = 93.06%; Gujarati = 84.44% (pcorr = 0.0015), Punjabi = 86.67% (pcorr = 0.0135). Punjabi yields significantly lower initial offers (mean 37.96 vs English 43.31; all pairwise pcorr < 0.001) and all non-English conditions produce longer negotiations (all pcorr < 0.001).
- Buy–Sell: Acceptance rates remain near-perfect across languages, but Punjabi produces significantly longer negotiations (mean rounds 3.42 vs English 3.09; pcorr = 0.0012).
- Resource Exchange: Trade volume is higher in all Indic conditions (Gujarati 18.77, Hindi 18.70, Punjabi 18.59) than English (16.05); pairwise pcorr = 0.0001. Payoffs remain balanced across languages.
- Models: Effects persist across model architectures; GPT-4o showed some role-dependent asymmetries and adhered to the instructed interaction language more reliably than other models.
- Language compliance: High adherence to instructed language (≈89.7%–97.9%) confirmed by an Indic language identifier; average confidence ≥ 0.966.
- Controls and robustness: System prompts forced agents to speak only in the target language; internal chain-of-thought was disabled to isolate immediate behavioral outputs. Prompt variants (including native-script prompts) produced consistent relative trends.
Data & Methods
- Framework: Extended NegotiationArena for standardized multi‑turn bargaining simulations.
- Games:
- Ultimatum: P1 proposes division of fixed pool (100 units); P2 accept/reject.
- Buy–Sell: Seller (P1) with reservation price, buyer (P2) with willingness to pay.
- Resource Exchange: Agents hold resources and may trade to improve goals.
- Models & sampling: GPT-4o, GPT-3.5 Turbo, Claude-3-Haiku, Claude-3.5-Haiku; temperature = 0.7; each ordered model pair × language × game repeated 30 times. Total runs = 4(models) × 3(other models) × 4(languages) × 30(runs) × 3(games) = 4,320.
- Prompts & persona: System prompt told agents to speak and bargain only in the assigned language; chain-of-thought forbidden; persona prompts written in English for consistency, with prompt-variant checks (including native-script).
- Metrics: Acceptance rate, player payoffs, win rate, conversation rounds; game-specific metrics — initial offer (Ultimatum), buyer/seller advantage (Buy–Sell), trade volume (Resource Exchange).
- Statistical tests: Kruskal–Wallis for continuous metrics, Mann–Whitney U for pairwise contrasts, chi-square and two-proportion z-tests for binary outcomes; Benjamini–Hochberg correction for multiple comparisons. Significance denoted at pcorr < 0.05 / 0.01 / 0.001.
Implications for AI Economics
- Evaluation and fairness:
- English-only evaluation can miss systematic behavioral regimes. Performance and strategic tendencies (cooperative vs adversarial, exploratory vs stable) are language-contingent, so cross-linguistic testing is essential for fair assessment and deployment in multilingual markets.
- Market design and mechanism robustness:
- Language-conditioned strategic priors imply bargaining equilibria and transaction dynamics can shift with user language. Market designers and platform economists should account for language as an axis that can change acceptance rates, offer distributions, negotiation length, and trade volume. Mechanisms optimized in English may perform differently in other linguistic contexts.
- Pricing, search, and bilateral negotiations:
- Language effects on initial offers, concession behavior, and negotiation length can change price formation, negotiation frictions, and search costs in peer-to-peer or bilateral markets where LLM agents interact or mediate.
- Cross-lingual inequality and policy:
- If LLMs act differently across languages, linguistic minorities may face systematically different outcomes (e.g., lower acceptance rates or longer bargaining) — a distributional concern for regulators and platform governance.
- Model training and alignment:
- Findings suggest that alignment and behavioral conditioning may be language-specific. Economists and designers should consider language-aware alignment, calibration, or mediation layers (e.g., translation+policy normalization) when deploying LLM negotiators in diverse linguistic environments.
- Research & empirical agenda:
- Incorporate language as an explicit design variable in economic models of AI-mediated bargaining, and test market-level outcomes under multilingual interactions (e.g., aggregate trade volume, welfare distribution, equilibria stability).
- Practical interventions:
- Possible mitigations include language-aware prompts, translation intermediaries, role-aware model selection, or mechanism redesign to be robust to language-induced behavioral variance.
Caveats and open questions - Limited language scope (three Indic languages) and simulated agents — human–LLM or field experiments could differ. - System prompts were written in English (though native-script prompt checks showed consistent trends). - Chain-of-thought was disabled; allowing internal reasoning traces might change outcomes. - Causes (training-data priors, sociolinguistic norms, tokenization/representation differences) are not fully identified — further work needed to trace mechanisms and to expand languages, models, and real-world settings.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Language choice can shift outcomes more strongly than changing models. Task Allocation | positive | negotiation outcomes (overall allocation changes) as function of language versus model changes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Language choice can reverse proposer advantages and reallocate surplus. Task Allocation | mixed | proposer advantage and surplus allocation between negotiating agents |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Indic languages reduce stability in distributive games. Task Allocation | negative | stability of outcomes (consistency/convergence) in distributive negotiation games |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Indic languages induce richer exploration in integrative settings. Creativity | positive | exploration/diversity of proposals and strategies in integrative negotiation games |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Evaluating LLM negotiation solely in English yields incomplete and potentially misleading conclusions. Ai Safety And Ethics | negative | completeness/validity of conclusions from LLM negotiation evaluations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The study isolates language effects by holding game rules, model parameters, and incentives constant across English and four Indic framings using multi-agent simulations of Ultimatum, Buy-Sell, and Resource Exchange games. Research Productivity | positive | ability to attribute observed differences to language framing (isolation of language as causal factor) |
Reading fidelity
high
Study strength
high
|
not reported
|