0 cumulative citations
View corpus contextAllowing groups to delegate decisions to an LLM increased joint surplus in experimental bargaining, yet participants overwhelmingly preferred higher-control advice and frequently modified AI proposals — reducing realized gains. The welfare improvement stems not from better model capability but from removing the human filter: autonomous execution captured surplus that advisory and coaching modes lost through user overrides.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes. We present an online behavioral experiment (N=243) in which participants play three multi-tu rn bargaining games in groups of three. Each game, presented in randomized order, grants access to a single LLM assistance modality: proactive recommendations from an Advisor, reactive feedback from a Coach, or autonomous execution by a Delegate. All three modalitie s are powered by an LLM with super-human performance within this negotiation setting. On each turn, participants privately decide whe ther to act manually or use the AI modality available in that game. We document a preference-performance misalignment: participants s trongly prefer the higher-control Advisor (44%) over the Delegate (19%), yet groups only significantly increase collective surplus un der Delegate access. Adjusting for voluntary non-compliance, delegating to the AI yields suggestive individual welfare gains, roughly 1.5x the intent-to-treat estimate. A mechanism analysis traces this gap to a human filter: AI-generated proposals create more joint surplus than manual proposals across all conditions, but in the Advisor and Coach modes users modify, override, or ignore the AI's su ggestions, reverting toward human-baseline trade patterns. The Delegate advantage arises not from a different AI capability but from bypassing this filtering step altogether. Realizing these welfare gains depends not only on model capability, but on the interaction structure through which that capability is delivered. We argue that assistance modalities should be designed as mechanisms with endog enous participation; adoption-compatible interaction rules are a prerequisite to improving welfare with automated assistance.
Summary
Main Finding
In a randomized, within-subjects lab experiment (N = 243), giving groups of three access to an autonomous Delegate (LLM executes trades) produced the largest increases in collective surplus, despite participants strongly preferring the higher-control Advisor (proactive recommendations). The welfare advantage of Delegation arises not from superior model capability but from bypassing a “human filter”: when users can edit or reject AI suggestions (Advisor/Coach), they systematically dilute high-quality AI proposals back toward human baseline behavior, undoing much of the potential welfare gain. Delegates both boost group surplus and generate positive externalities for unassisted counterparts.
Key Points
- Interaction modalities tested (same underlying LLM capability, Gemini-2.5 Flash prompt scaffolding):
- Advisor: AI proactively proposes offers; human has veto/edit power.
- Coach: Human proposes; AI provides feedback before execution.
- Delegate: AI autonomously executes proposals on behalf of the human.
- Experimental design:
- N = 243 participants (81 three-person groups).
- Within-subjects: each participant played three bargaining games (one per modality) in randomized order.
- On each turn participants privately chose to act manually or use the available AI modality.
- Preference–performance misalignment:
- Participants strongly preferred Advisor-type control (reported preference ~44%) over Delegate (~19%).
- Despite this, only Delegate access produced a statistically significant increase in group surplus relative to human baseline.
- Adoption and treatment compliance:
- Voluntary non-compliance (users often ignored/edited AI suggestions) reduced realized treatment effects in Advisor/Coach arms.
- After adjusting for voluntary non-compliance, delegation yields suggestive individual welfare gains approximately 1.5× the intent-to-treat (ITT) estimate.
- Mechanism — the “human filter”:
- AI-generated proposals were higher-quality (created more joint surplus) than manual proposals across all modalities.
- In Advisor and Coach modes, users frequently modified, overrode, or ignored AI proposals; these edits reverted offers toward human norms (e.g., fairness/1-for-1 trades), reducing welfare gains.
- Delegate bypasses this filter, preserving AI proposal quality and its surplus benefits.
- Market-making spillovers:
- Delegate adoption produced positive externalities: unassisted counterparts’ surplus increased (reported ~21.6% uplift), consistent with delegates acting as cooperative market-makers rather than exploitative arbitrageurs.
- Benchmarking:
- Human-only baseline (from prior work, same game mechanics): mean scaled group surplus ≈ 0.537 (±0.024). LLM agents exceed human-only performance in All-Agent baselines, isolating interaction-structure effects.
Data & Methods
- Game environment:
- Stylized chip-exchange bargaining game (Qian et al. 2025b). Each game: nine turns; on a proposing turn one player suggests a trade; two other players simultaneously accept/decline; if both accept, one trade is randomly executed. Players have randomly assigned private chip valuations; objective is to maximize surplus relative to Pareto optimum.
- Strategic interdependence: payoff coupling, opportunity-set and informational externalities across turns and players.
- Treatments:
- Group-level randomization to one of the three assistance modalities per game; within-subjects counterbalanced order so each participant experienced all three modalities across games.
- AI:
- LLM-based agents (Gemini-2.5 Flash scaffolding) tuned to be superhuman in this negotiation setting; identical model capability across modalities to isolate interaction-structure effects.
- Outcome measures:
- Scaled group surplus and scaled individual surplus (surplus achieved divided by the Pareto-efficient maximum for that game).
- Intent-to-Treat (ITT) effects and compliance-adjusted estimates (to account for voluntary non-use or modification of AI suggestions).
- Mechanism analysis:
- Comparison of AI-generated proposals vs manual proposals on realized joint surplus.
- Tracked user edits/overrides/ignores of AI suggestions to quantify the “human filter.”
- Measured spillovers to unassisted players in mixed-adoption scenarios.
Implications for AI Economics
- Interaction modality matters as much as model capability:
- Economic gains from AI in multi-agent settings depend on the allocation of decision initiative; providing powerful agents is not sufficient if users systematically intervene in ways that undo gains.
- Endogenous adoption and equilibrium effects:
- Adoption is endogenous and shaped by preferences for control; researchers and policymakers should measure both ITT and compliance-adjusted (LATE-like) effects when evaluating welfare impacts.
- Strategic externalities mean an individual’s delegation choice alters others’ opportunity sets and welfare—standard 1:1 human-AI evaluations can miss important systemic effects.
- Positive market-making role for autonomous agents:
- Delegates can raise collective surplus and produce positive spillovers (21.6% uplift for unassisted counterparts in this study). This suggests potential for AI agents to structurally improve market or negotiation efficiency when allowed to act autonomously.
- Design and mechanism considerations:
- Interface and institutional rules (who has initiative, veto rights, timing of confidence disclosure) are mechanism design parameters that can make or break welfare gains.
- Adoption-compatible interaction rules (e.g., limited veto windows, progressive confidence signals, group-level performance metrics) may be necessary to realize aggregate gains while preserving perceived control and accountability.
- Policy and organizational implications:
- If welfare gains from delegation generalize, firms and platforms might obtain aggregate productivity gains by enabling controlled delegation regimes; however, distributional consequences, accountability, and norm effects must be assessed.
- Regulatory assessments should account for equilibrium and spillover effects (both positive and negative) of agentic technologies in multi-agent contexts.
- Generalizability caveat:
- Results are from a stylized bargaining game with induced valuations and laboratory incentives. Whether the same preference–performance misalignment and human-filter dynamics hold in other strategic domains, longer time horizons, richer communication channels, or real-world institutional environments remains an open empirical question.
Summary takeaway: In multi-party strategic environments, the structure of human–AI interaction—who initiates and who can veto—shapes adoption and collective welfare as much as the underlying AI capability. Mechanism design that aligns adoption incentives with welfare (rather than only increasing model quality) is central to realizing the economic benefits of agentic AIs.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Participants strongly prefer the higher-control Advisor (44%) over the Delegate (19%). Adoption Rate | positive | participant preference for assistance modality |
Reading fidelity
high
Study strength
high
|
n=243
44% vs 19%
|
| Groups only significantly increase collective surplus under Delegate access (i.e., groups significantly increase collective surplus when using the Delegate but not under Advisor or Coach). Team Performance | positive | collective surplus (group-level payoff/welfare) |
Reading fidelity
high
Study strength
high
|
n=243
|
| Adjusting for voluntary non-compliance, delegating to the AI yields suggestive individual welfare gains roughly 1.5x the intent-to-treat estimate. Wages | positive | individual welfare (participant payoff) |
Reading fidelity
high
Study strength
medium
|
n=243
1.5x
|
| AI-generated proposals create more joint surplus than manual proposals across all conditions. Team Performance | positive | joint surplus produced by proposals |
Reading fidelity
high
Study strength
medium
|
n=243
|
| In the Advisor and Coach modes, users modify, override, or ignore the AI's suggestions, reverting toward human-baseline trade patterns (a 'human filter' that reduces AI-generated surplus). Decision Quality | negative | degree of user modification/override of AI suggestions and resulting deviation from AI-optimal proposals |
Reading fidelity
high
Study strength
medium
|
n=243
|
| The Delegate advantage arises not from a different AI capability but from bypassing the human filtering step altogether. Team Performance | positive | realized group welfare (collective surplus) attributable to interaction structure rather than model capability |
Reading fidelity
high
Study strength
medium
|
n=243
|
| All three modalities (Advisor, Coach, Delegate) are powered by an LLM with super-human performance within this negotiation setting. Other | positive | LLM performance on negotiation tasks (benchmark vs human baseline) |
Reading fidelity
high
Study strength
medium
|
not reported
|