The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Allowing groups to delegate decisions to an LLM increased joint surplus in experimental bargaining, yet participants overwhelmingly preferred higher-control advice and frequently modified AI proposals — reducing realized gains. The welfare improvement stems not from better model capability but from removing the human filter: autonomous execution captured surplus that advisory and coaching modes lost through user overrides.

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation
Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian · February 12, 2026
arxiv rct medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Kehang Zhu unresolved corpus identity
  2. Nithum Thain unresolved corpus identity
  3. Vivian Tsai unresolved corpus identity
  4. James Wexler unresolved corpus identity
  5. Crystal Qian unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Kehang Zhu provider ID
  2. Nithum Thain provider ID
  3. Vivian Tsai provider ID
  4. James Wexler provider ID
  5. Crystal Qian provider ID
In three-player bargaining games, autonomous delegation to an LLM raised collective surplus, but participants preferred higher-control assistance and often filtered or ignored AI suggestions in advisory modes, undoing much of the AI's potential benefit.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes. We present an online behavioral experiment (N=243) in which participants play three multi-tu rn bargaining games in groups of three. Each game, presented in randomized order, grants access to a single LLM assistance modality: proactive recommendations from an Advisor, reactive feedback from a Coach, or autonomous execution by a Delegate. All three modalitie s are powered by an LLM with super-human performance within this negotiation setting. On each turn, participants privately decide whe ther to act manually or use the AI modality available in that game. We document a preference-performance misalignment: participants s trongly prefer the higher-control Advisor (44%) over the Delegate (19%), yet groups only significantly increase collective surplus un der Delegate access. Adjusting for voluntary non-compliance, delegating to the AI yields suggestive individual welfare gains, roughly 1.5x the intent-to-treat estimate. A mechanism analysis traces this gap to a human filter: AI-generated proposals create more joint surplus than manual proposals across all conditions, but in the Advisor and Coach modes users modify, override, or ignore the AI's su ggestions, reverting toward human-baseline trade patterns. The Delegate advantage arises not from a different AI capability but from bypassing this filtering step altogether. Realizing these welfare gains depends not only on model capability, but on the interaction structure through which that capability is delivered. We argue that assistance modalities should be designed as mechanisms with endog enous participation; adoption-compatible interaction rules are a prerequisite to improving welfare with automated assistance.

Summary

Main Finding

In a randomized, within-subjects lab experiment (N = 243), giving groups of three access to an autonomous Delegate (LLM executes trades) produced the largest increases in collective surplus, despite participants strongly preferring the higher-control Advisor (proactive recommendations). The welfare advantage of Delegation arises not from superior model capability but from bypassing a “human filter”: when users can edit or reject AI suggestions (Advisor/Coach), they systematically dilute high-quality AI proposals back toward human baseline behavior, undoing much of the potential welfare gain. Delegates both boost group surplus and generate positive externalities for unassisted counterparts.

Key Points

  • Interaction modalities tested (same underlying LLM capability, Gemini-2.5 Flash prompt scaffolding):
    • Advisor: AI proactively proposes offers; human has veto/edit power.
    • Coach: Human proposes; AI provides feedback before execution.
    • Delegate: AI autonomously executes proposals on behalf of the human.
  • Experimental design:
    • N = 243 participants (81 three-person groups).
    • Within-subjects: each participant played three bargaining games (one per modality) in randomized order.
    • On each turn participants privately chose to act manually or use the available AI modality.
  • Preference–performance misalignment:
    • Participants strongly preferred Advisor-type control (reported preference ~44%) over Delegate (~19%).
    • Despite this, only Delegate access produced a statistically significant increase in group surplus relative to human baseline.
  • Adoption and treatment compliance:
    • Voluntary non-compliance (users often ignored/edited AI suggestions) reduced realized treatment effects in Advisor/Coach arms.
    • After adjusting for voluntary non-compliance, delegation yields suggestive individual welfare gains approximately 1.5× the intent-to-treat (ITT) estimate.
  • Mechanism — the “human filter”:
    • AI-generated proposals were higher-quality (created more joint surplus) than manual proposals across all modalities.
    • In Advisor and Coach modes, users frequently modified, overrode, or ignored AI proposals; these edits reverted offers toward human norms (e.g., fairness/1-for-1 trades), reducing welfare gains.
    • Delegate bypasses this filter, preserving AI proposal quality and its surplus benefits.
  • Market-making spillovers:
    • Delegate adoption produced positive externalities: unassisted counterparts’ surplus increased (reported ~21.6% uplift), consistent with delegates acting as cooperative market-makers rather than exploitative arbitrageurs.
  • Benchmarking:
    • Human-only baseline (from prior work, same game mechanics): mean scaled group surplus ≈ 0.537 (±0.024). LLM agents exceed human-only performance in All-Agent baselines, isolating interaction-structure effects.

Data & Methods

  • Game environment:
    • Stylized chip-exchange bargaining game (Qian et al. 2025b). Each game: nine turns; on a proposing turn one player suggests a trade; two other players simultaneously accept/decline; if both accept, one trade is randomly executed. Players have randomly assigned private chip valuations; objective is to maximize surplus relative to Pareto optimum.
    • Strategic interdependence: payoff coupling, opportunity-set and informational externalities across turns and players.
  • Treatments:
    • Group-level randomization to one of the three assistance modalities per game; within-subjects counterbalanced order so each participant experienced all three modalities across games.
  • AI:
    • LLM-based agents (Gemini-2.5 Flash scaffolding) tuned to be superhuman in this negotiation setting; identical model capability across modalities to isolate interaction-structure effects.
  • Outcome measures:
    • Scaled group surplus and scaled individual surplus (surplus achieved divided by the Pareto-efficient maximum for that game).
    • Intent-to-Treat (ITT) effects and compliance-adjusted estimates (to account for voluntary non-use or modification of AI suggestions).
  • Mechanism analysis:
    • Comparison of AI-generated proposals vs manual proposals on realized joint surplus.
    • Tracked user edits/overrides/ignores of AI suggestions to quantify the “human filter.”
    • Measured spillovers to unassisted players in mixed-adoption scenarios.

Implications for AI Economics

  • Interaction modality matters as much as model capability:
    • Economic gains from AI in multi-agent settings depend on the allocation of decision initiative; providing powerful agents is not sufficient if users systematically intervene in ways that undo gains.
  • Endogenous adoption and equilibrium effects:
    • Adoption is endogenous and shaped by preferences for control; researchers and policymakers should measure both ITT and compliance-adjusted (LATE-like) effects when evaluating welfare impacts.
    • Strategic externalities mean an individual’s delegation choice alters others’ opportunity sets and welfare—standard 1:1 human-AI evaluations can miss important systemic effects.
  • Positive market-making role for autonomous agents:
    • Delegates can raise collective surplus and produce positive spillovers (21.6% uplift for unassisted counterparts in this study). This suggests potential for AI agents to structurally improve market or negotiation efficiency when allowed to act autonomously.
  • Design and mechanism considerations:
    • Interface and institutional rules (who has initiative, veto rights, timing of confidence disclosure) are mechanism design parameters that can make or break welfare gains.
    • Adoption-compatible interaction rules (e.g., limited veto windows, progressive confidence signals, group-level performance metrics) may be necessary to realize aggregate gains while preserving perceived control and accountability.
  • Policy and organizational implications:
    • If welfare gains from delegation generalize, firms and platforms might obtain aggregate productivity gains by enabling controlled delegation regimes; however, distributional consequences, accountability, and norm effects must be assessed.
    • Regulatory assessments should account for equilibrium and spillover effects (both positive and negative) of agentic technologies in multi-agent contexts.
  • Generalizability caveat:
    • Results are from a stylized bargaining game with induced valuations and laboratory incentives. Whether the same preference–performance misalignment and human-filter dynamics hold in other strategic domains, longer time horizons, richer communication channels, or real-world institutional environments remains an open empirical question.

Summary takeaway: In multi-party strategic environments, the structure of human–AI interaction—who initiates and who can veto—shapes adoption and collective welfare as much as the underlying AI capability. Mechanism design that aligns adoption incentives with welfare (rather than only increasing model quality) is central to realizing the economic benefits of agentic AIs.

Assessment

Paper Typerct Evidence Strengthmedium — Strong internal identification from randomized assignment and ITT/complier adjustments yields credible causal claims about how interaction modality affects outcomes in the experimental task; however, external validity is limited by a single lab-style online bargaining task, modest sample size (N=243), a constrained three-player setting, and use of one tuned LLM, so generalization to real-world organizational or market settings is uncertain. Methods Rigorhigh — Well-powered randomized design with randomized order, clear ITT estimands, adjustment for non-compliance (instrumental/complier analysis), and a mechanism analysis comparing AI vs manual proposals and user filtering behavior; pre-registered hypotheses or robustness checks are not mentioned here, but core randomization and analytical steps are appropriate and carefully targeted to the causal question. SampleOnline behavioral experiment with N=243 participants grouped into three-person teams; each participant played three multi-turn bargaining games (one per assigned LLM modality: Advisor, Coach, Delegate) presented in randomized order; participants could choose on each turn whether to act manually or use the available AI; AI models were LLMs tuned to deliver super-human negotiation suggestions/executions within the experimental context. Themeshuman_ai_collab productivity adoption org_design IdentificationWithin-subject randomized experiment assigning each three-person bargaining game to one of three LLM assistance modalities (Advisor, Coach, Delegate) with randomized order; primary comparisons are intent-to-treat (ITT) differences in collective surplus by assigned modality, with additional complier-adjusted estimates that instrument actual AI usage by assignment to adjust for voluntary non-compliance; analyses control for game/order effects and use within-subject variation to isolate causal effects of interaction modality. GeneralizabilityArtificial lab-style bargaining games may not reflect complexity of real workplace or market negotiations, Likely WEIRD/online subject pool (limits demographic representativeness), Single LLM and specific prompt/tuning used — results may differ with other models or domains, Short-run interactions; long-run learning, strategic responses, and repeated-market dynamics not observed, Three-player groups only — scaling to larger teams or organizational structures unclear

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Participants strongly prefer the higher-control Advisor (44%) over the Delegate (19%). Adoption Rate positive participant preference for assistance modality
Reading fidelity high
Study strength high
n=243
44% vs 19%
1.0
Groups only significantly increase collective surplus under Delegate access (i.e., groups significantly increase collective surplus when using the Delegate but not under Advisor or Coach). Team Performance positive collective surplus (group-level payoff/welfare)
Reading fidelity high
Study strength high
n=243
1.0
Adjusting for voluntary non-compliance, delegating to the AI yields suggestive individual welfare gains roughly 1.5x the intent-to-treat estimate. Wages positive individual welfare (participant payoff)
Reading fidelity high
Study strength medium
n=243
1.5x
0.6
AI-generated proposals create more joint surplus than manual proposals across all conditions. Team Performance positive joint surplus produced by proposals
Reading fidelity high
Study strength medium
n=243
0.6
In the Advisor and Coach modes, users modify, override, or ignore the AI's suggestions, reverting toward human-baseline trade patterns (a 'human filter' that reduces AI-generated surplus). Decision Quality negative degree of user modification/override of AI suggestions and resulting deviation from AI-optimal proposals
Reading fidelity high
Study strength medium
n=243
0.6
The Delegate advantage arises not from a different AI capability but from bypassing the human filtering step altogether. Team Performance positive realized group welfare (collective surplus) attributable to interaction structure rather than model capability
Reading fidelity high
Study strength medium
n=243
0.6
All three modalities (Advisor, Coach, Delegate) are powered by an LLM with super-human performance within this negotiation setting. Other positive LLM performance on negotiation tasks (benchmark vs human baseline)
Reading fidelity high
Study strength medium
not reported
0.6

Notes