The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Language alters LLM negotiation outcomes more than model changes: switching from English to several Indic languages can reverse proposer advantages and shift surplus allocation; distributive games become less stable while integrative games induce more exploratory deals.

The Language of Bargaining: Linguistic Effects in LLM Negotiations
Stuti Sinha, Himanshu Kumar, Aryan Raju Mandapati, Rakshit Sakhuja, Dhruv Kumar · January 07, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Stuti Sinha unresolved corpus identity
  2. Himanshu Kumar unresolved corpus identity
  3. Aryan Raju Mandapati unresolved corpus identity
  4. Rakshit Sakhuja unresolved corpus identity
  5. Dhruv Kumar unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Stuti Sinha provider ID
  2. Himanshu kumar provider ID
  3. Aryan Raju Mandapati provider ID
  4. Rakshit Sakhuja provider ID
  5. Dhruv Kumar provider ID
When identical LLM negotiators operate in different languages, language choice substantially alters negotiation outcomes—sometimes more than swapping models—reversing proposer advantage and reallocating surplus, with effects varying by game type.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Negotiation is a core component of social intelligence, requiring agents to balance strategic reasoning, cooperation, and social norms. Recent work shows that LLMs can engage in multi-turn negotiation, yet nearly all evaluations occur exclusively in English. Using controlled multi-agent simulations across Ultimatum, Buy-Sell, and Resource Exchange games, we systematically isolate language effects across English and four Indic framings (Hindi, Punjabi, Gujarati, Marwadi) by holding game rules, model parameters, and incentives constant across all conditions. We find that language choice can shift outcomes more strongly than changing models, reversing proposer advantages and reallocating surplus. Crucially, effects are task-contingent: Indic languages reduce stability in distributive games yet induce richer exploration in integrative settings. Our results demonstrate that evaluating LLM negotiation solely in English yields incomplete and potentially misleading conclusions. These findings caution against English-only evaluation of LLMs and suggest that culturally-aware evaluation is essential for fair deployment.

Summary

Main Finding

The interaction language systematically alters LLM negotiation behavior. Holding game rules, incentives, and model parameters constant across 4,320 simulated games, the authors show language choice acts as a strategic prior that changes outcomes and dynamics in task-dependent ways: Indic framings (Hindi, Gujarati, Punjabi) reduce stability in distributive games (Ultimatum, Buy–Sell) — longer negotiations and lower acceptance in some cases — while increasing exploratory trade (higher trade volumes) in an integrative Resource Exchange game. These effects persist across multiple LLM architectures and are robust to prompt variants and script choice.

Key Points

  • Experiments: 4,320 multi-agent games covering three canonical negotiation settings (Ultimatum, Buy–Sell, Resource Exchange), four language conditions (English baseline; Hindi, Gujarati, Punjabi), and four multilingual LLMs (GPT-4o, GPT-3.5 Turbo, Claude-3-Haiku, Claude-3.5-Haiku).
  • Statistically significant language effects (nonparametric tests; Benjamini–Hochberg correction):
    • Ultimatum: English acceptance rate = 93.06%; Gujarati = 84.44% (pcorr = 0.0015), Punjabi = 86.67% (pcorr = 0.0135). Punjabi yields significantly lower initial offers (mean 37.96 vs English 43.31; all pairwise pcorr < 0.001) and all non-English conditions produce longer negotiations (all pcorr < 0.001).
    • Buy–Sell: Acceptance rates remain near-perfect across languages, but Punjabi produces significantly longer negotiations (mean rounds 3.42 vs English 3.09; pcorr = 0.0012).
    • Resource Exchange: Trade volume is higher in all Indic conditions (Gujarati 18.77, Hindi 18.70, Punjabi 18.59) than English (16.05); pairwise pcorr = 0.0001. Payoffs remain balanced across languages.
  • Models: Effects persist across model architectures; GPT-4o showed some role-dependent asymmetries and adhered to the instructed interaction language more reliably than other models.
  • Language compliance: High adherence to instructed language (≈89.7%–97.9%) confirmed by an Indic language identifier; average confidence ≥ 0.966.
  • Controls and robustness: System prompts forced agents to speak only in the target language; internal chain-of-thought was disabled to isolate immediate behavioral outputs. Prompt variants (including native-script prompts) produced consistent relative trends.

Data & Methods

  • Framework: Extended NegotiationArena for standardized multi‑turn bargaining simulations.
  • Games:
    • Ultimatum: P1 proposes division of fixed pool (100 units); P2 accept/reject.
    • Buy–Sell: Seller (P1) with reservation price, buyer (P2) with willingness to pay.
    • Resource Exchange: Agents hold resources and may trade to improve goals.
  • Models & sampling: GPT-4o, GPT-3.5 Turbo, Claude-3-Haiku, Claude-3.5-Haiku; temperature = 0.7; each ordered model pair × language × game repeated 30 times. Total runs = 4(models) × 3(other models) × 4(languages) × 30(runs) × 3(games) = 4,320.
  • Prompts & persona: System prompt told agents to speak and bargain only in the assigned language; chain-of-thought forbidden; persona prompts written in English for consistency, with prompt-variant checks (including native-script).
  • Metrics: Acceptance rate, player payoffs, win rate, conversation rounds; game-specific metrics — initial offer (Ultimatum), buyer/seller advantage (Buy–Sell), trade volume (Resource Exchange).
  • Statistical tests: Kruskal–Wallis for continuous metrics, Mann–Whitney U for pairwise contrasts, chi-square and two-proportion z-tests for binary outcomes; Benjamini–Hochberg correction for multiple comparisons. Significance denoted at pcorr < 0.05 / 0.01 / 0.001.

Implications for AI Economics

  • Evaluation and fairness:
    • English-only evaluation can miss systematic behavioral regimes. Performance and strategic tendencies (cooperative vs adversarial, exploratory vs stable) are language-contingent, so cross-linguistic testing is essential for fair assessment and deployment in multilingual markets.
  • Market design and mechanism robustness:
    • Language-conditioned strategic priors imply bargaining equilibria and transaction dynamics can shift with user language. Market designers and platform economists should account for language as an axis that can change acceptance rates, offer distributions, negotiation length, and trade volume. Mechanisms optimized in English may perform differently in other linguistic contexts.
  • Pricing, search, and bilateral negotiations:
    • Language effects on initial offers, concession behavior, and negotiation length can change price formation, negotiation frictions, and search costs in peer-to-peer or bilateral markets where LLM agents interact or mediate.
  • Cross-lingual inequality and policy:
    • If LLMs act differently across languages, linguistic minorities may face systematically different outcomes (e.g., lower acceptance rates or longer bargaining) — a distributional concern for regulators and platform governance.
  • Model training and alignment:
    • Findings suggest that alignment and behavioral conditioning may be language-specific. Economists and designers should consider language-aware alignment, calibration, or mediation layers (e.g., translation+policy normalization) when deploying LLM negotiators in diverse linguistic environments.
  • Research & empirical agenda:
    • Incorporate language as an explicit design variable in economic models of AI-mediated bargaining, and test market-level outcomes under multilingual interactions (e.g., aggregate trade volume, welfare distribution, equilibria stability).
  • Practical interventions:
    • Possible mitigations include language-aware prompts, translation intermediaries, role-aware model selection, or mechanism redesign to be robust to language-induced behavioral variance.

Caveats and open questions - Limited language scope (three Indic languages) and simulated agents — human–LLM or field experiments could differ. - System prompts were written in English (though native-script prompt checks showed consistent trends). - Chain-of-thought was disabled; allowing internal reasoning traces might change outcomes. - Causes (training-data priors, sociolinguistic norms, tokenization/representation differences) are not fully identified — further work needed to trace mechanisms and to expand languages, models, and real-world settings.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Within-simulation variation isolates language as the treatment and shows consistent, task-dependent effects, supporting a credible causal link in the experimental setting; however, evidence is limited to simulated LLM agents, a limited set of languages, games, and (apparently) model variants, so external validity to human-AI interactions and real-world markets is uncertain. Methods Rigormedium — The study systematically controls key variables and compares multiple games and languages, improving internal validity; but rigor depends on details not provided here (number and diversity of model instances, seed/sample sizes, prompt standardization, translation quality checks, robustness tests across model families), so methodological thoroughness appears solid but not exhaustive. SampleSimulated multi-agent interactions between LLM-based negotiators across three game families (Ultimatum, Buy-Sell, Resource Exchange), run in English and four Indic languages (Hindi, Punjabi, Gujarati, Marwadi), with game rules, incentives, and model parameters held constant across language conditions; outcomes aggregated over repeated simulation runs. Themeshuman_ai_collab inequality IdentificationControlled multi-agent simulations that hold game rules, model parameters, prompts, and incentives constant while varying only the language (English vs four Indic languages) across repeated runs of Ultimatum, Buy-Sell, and Resource Exchange games to attribute observed differences in outcomes to language choice. GeneralizabilityResults are from simulated LLM agents, not human–AI or human–human negotiations, limiting external validity to real markets and workplaces., Only five languages were tested (English + four Indic languages), so findings may not generalize to other languages or dialects., Effects may depend on the particular LLM family, size, or training data; limited model diversity would constrain generalizability., Game types are stylized (Ultimatum, Buy-Sell, Resource Exchange) and may not capture complexity of real-world negotiations., Potential confounds like translation/terminology quality, cultural framing in prompts, or tokenization differences may mediate language effects.

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Language choice can shift outcomes more strongly than changing models. Task Allocation positive negotiation outcomes (overall allocation changes) as function of language versus model changes
Reading fidelity high
Study strength medium
not reported
0.48
Language choice can reverse proposer advantages and reallocate surplus. Task Allocation mixed proposer advantage and surplus allocation between negotiating agents
Reading fidelity high
Study strength medium
not reported
0.48
Indic languages reduce stability in distributive games. Task Allocation negative stability of outcomes (consistency/convergence) in distributive negotiation games
Reading fidelity high
Study strength medium
not reported
0.48
Indic languages induce richer exploration in integrative settings. Creativity positive exploration/diversity of proposals and strategies in integrative negotiation games
Reading fidelity high
Study strength medium
not reported
0.48
Evaluating LLM negotiation solely in English yields incomplete and potentially misleading conclusions. Ai Safety And Ethics negative completeness/validity of conclusions from LLM negotiation evaluations
Reading fidelity high
Study strength medium
not reported
0.48
The study isolates language effects by holding game rules, model parameters, and incentives constant across English and four Indic framings using multi-agent simulations of Ultimatum, Buy-Sell, and Resource Exchange games. Research Productivity positive ability to attribute observed differences to language framing (isolation of language as causal factor)
Reading fidelity high
Study strength high
not reported
0.8

Notes