The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

People judge machine-made offers as less socially appropriate and are more willing to reject them, yet accept machine-enforced rejections as no less appropriate than human ones; in short, machines are held to different fairness norms for decisions but not for enforcement.

Do people expect different behavior from large language models acting on their behalf? Evidence from norm elicitations in two canonical economic games
Paweł Niszczota, Elia Antoniou · January 14, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Paweł Niszczota unresolved corpus identity
  2. Elia Antoniou unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Pawel Niszczota provider ID
  2. E. Antoniou provider ID
In incentivized, pre-registered experiments with representative US and UK samples, people rate identical resource offers made by LLMs as less socially appropriate than human-made offers, are more inclined to see rejecting machine-made offers as appropriate, yet view machine rejections as no less appropriate than human rejections.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

While delegating tasks to large language models (LLMs) can save people time, there is growing evidence that offloading tasks to such models produces social costs. We use behavior in two canonical economic games to study whether people have different expectations when decisions are made by LLMs acting on their behalf instead of themselves. More specifically, we study the social appropriateness of a spectrum of possible behaviors: when LLMs divide resources on our behalf (Dictator Game and Ultimatum Game) and when they monitor the fairness of splits of resources (Ultimatum Game). We use the Krupka-Weber norm elicitation task to detect shifts in social appropriateness ratings. Results of two pre-registered and incentivized experimental studies using representative samples from the UK and US (N = 2,658) show three key findings. First, people find that offers from machines - when no acceptance is necessary - are judged to be less appropriate than when they come from humans, although there is no shift in the modal response. Second - when acceptance is necessary - it is more appropriate for a person to reject offers from machines than from humans. Third, receiving a rejection of an offer from a machine is no less socially appropriate than receiving the same rejection from a human. Overall, these results suggest that people apply different norms for machines deciding on how to split resources but are not opposed to machines enforcing the norms. The findings are consistent with offers made by machines now being viewed as having both a cognitive and emotional component.

Summary

Main Finding

People apply different social norms to LLMs acting on their behalf versus humans: offers generated by LLMs (when no acceptance is required) are judged less socially appropriate than identical offers from humans, whereas people are more willing to have (or expect) humans to reject offers made by LLMs; receiving rejections issued by LLMs is judged no less appropriate than receiving the same rejection from a human. Overall, norms differ for machines as proposers but not for machines as enforcers.

Key Points

  • Study scope: Two pre-registered, incentivized experiments using Krupka–Weber norm elicitation on representative UK and US samples (combined N = 2,658). Data and materials: https://osf.io/2yt9e/overview?view_only=8268c7ac95d94f2390fd3ed6dc76269a.
  • Games used:
    • Dictator Game: measures perceived appropriateness of 11 possible splits when the decision-maker is either a human or an LLM acting on a human’s behalf.
    • Ultimatum Game: elicit norms about rejecting offers and about who (human vs LLM) issues rejections.
  • Core empirical findings (from the two experiments):
  • In the Dictator Game (no acceptance mechanism), identical offers from LLMs are rated as less socially appropriate than identical offers from humans. The modal (most common) acceptable split did not shift, but appropriateness ratings decreased for machine-originated offers.
  • In the Ultimatum Game where acceptance/rejection matters, it is judged more appropriate for a person to reject offers made by LLMs than to reject identical offers made by humans.
  • A rejection coming from an LLM is judged no less socially appropriate than the same rejection coming from a human — i.e., people do not penalize machines for enforcing norms via rejection.
  • Interpretation: People appear to hold LLM proposers to different normative standards (or view machine-originated offers as carrying a different social meaning, possibly more exploitative), while accepting machines as legitimate enforcers of fairness.
  • Robustness & preregistration: analyses used incentive-compatible modal-matching, nonparametric tests (clustered bootstrap, Wilcoxon rank-sum), Holm–Bonferroni correction for multiple comparisons; study preregistered and ethics-approved.

Data & Methods

  • Design:
    • Mixed design in Experiment 1: each participant rated all 11 possible splits; between-subjects manipulation of agent type (Human vs LLM acting on behalf of a human).
    • Experiment 2 similarly tested norms about rejection behavior in the Ultimatum Game and legitimacy of LLMs as enforcers.
  • Norm elicitation: Krupka–Weber procedure — participants earn a bonus by matching the modal response in their condition, which incentivizes truthful reporting of perceived social norms.
  • Samples:
    • Representative samples from Prolific for the UK and a representative US sample (combined N = 2,658). Experiment 1: UK sample N ≈ 1,153 after exclusions (576 Human, 577 LLM).
  • Incentives: base participation fee + bonus (£0.50) for matching modal responses in a randomly selected choice (in Experiment 1).
  • Statistical approach:
    • Nonparametric tests to compare appropriateness ratings across Human vs LLM conditions.
    • Clustered bootstrap tests for shifts in modal peak; clustered Wilcoxon rank-sum tests for overall differences; Holm–Bonferroni correction for multiple split-wise tests.
  • Transparency: pre-registration, full data and code available on OSF; ethical approval obtained.

Implications for AI Economics

  • Delegation & social equilibria: Widespread delegation of distributive decisions to LLMs could alter norm perceptions and therefore aggregate outcomes (e.g., lower offers accepted socially when machines decide), potentially changing market or institutional equilibria.
  • Strategic use of AI and blame-shifting: People may view low offers from LLMs as more exploitative or as a deliberate strategy to avoid human blame; marketplace actors could exploit or be penalized by these perceptions.
  • Legitimacy of algorithmic enforcement: LLMs are not penalized for enforcing norms (e.g., rejecting unfair offers). This implies scope for automated enforcement mechanisms (moderation, sanctions, contract-termination agents) that the public may accept.
  • Design & deployment guidance:
    • Systems intended to propose allocations should be designed with awareness of negative normative judgments (e.g., favoring more egalitarian or “socially acceptable” proposals).
    • Systems used for enforcement can be effective socially, but transparency about role (proposer vs enforcer) matters.
  • Policy considerations:
    • Regulation and governance should distinguish between roles AIs play (decision-maker vs enforcer) rather than apply blanket approvals or bans.
    • Public attitudes are nuanced: policymakers should avoid one-size-fits-all restrictions and instead target role-specific guidelines (e.g., limits or disclosure requirements for LLMs making distributive proposals).
  • Research directions:
    • Link norm-elicited beliefs to actual delegation behavior and downstream market outcomes.
    • Explore heterogeneity (trust in AI, frequency of AI use, cultural differences) and longer-term effects as LLMs become more familiar.
    • Test interventions (framing, disclosure, calibrated fairness objectives) to mitigate negative normative reactions to machine-originated proposals.

If useful, I can produce a one-page bulletized summary for nontechnical audiences, or extract key statistical test results and effect sizes from the Supplementary Materials.

Assessment

Paper Typerct Evidence Strengthmedium — Random assignment, pre-registration, incentivized payments, and large representative samples (N=2,658) give credible causal identification for how actor identity affects social-appropriateness judgments in these game contexts, but the outcomes are normative ratings in stylized games rather than observed economic behavior or long-run market outcomes, and external validity to real-world delegation is limited. Methods Rigorhigh — Study uses standard, validated elicitation (Krupka–Weber), pre-registration, incentivized monetary games (Dictator and Ultimatum), adequate sample size, and representative US/UK samples; potential concerns are acknowledged (framing, stake size, description of LLMs), but internal validity and execution appear rigorous. SampleTwo pre-registered, incentivized online experiments with combined N = 2,658 participants drawn to be representative of the UK and US adult populations; participants made/responded to monetary splits in Dictator and Ultimatum Game scenarios and provided Krupka–Weber appropriateness ratings under randomized 'machine' vs 'human' actor descriptions. Themeshuman_ai_collab governance IdentificationRandomized online experiments: participants in representative UK and US samples were randomly assigned to conditions where resource splits and rejections were described as made by either an LLM (machine) or a human; social appropriateness was measured using the incentivized Krupka–Weber norm elicitation task within Dictator and Ultimatum Game settings (pre-registered). GeneralizabilityResults are from stylized monetary games (Dictator/Ultimatum) and may not generalize to complex, real-world delegation tasks or workplace settings, Online experimental framing and small stakes may not capture emotional, legal, or long-term consequences of delegation, Findings are limited to UK and US cultural contexts and may not hold in other countries, Characterization of 'LLM' in the study may not match specific deployed systems or future LLM capabilities, Norms and perceptions may evolve quickly as public familiarity with AI changes

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Offers from machines - when no acceptance is necessary - are judged to be less appropriate than when they come from humans. Decision Quality negative social appropriateness ratings of offers (Krupka-Weber elicitation) when no acceptance is required
Reading fidelity high
Study strength high
n=2658
1.0
There is no shift in the modal response for appropriateness ratings even though offers from machines are judged less appropriate. Decision Quality null_result modal response of social appropriateness ratings
Reading fidelity high
Study strength medium
n=2658
0.6
When acceptance is necessary (Ultimatum Game), it is more appropriate for a person to reject offers from machines than from humans. Decision Quality positive appropriateness ratings of rejecting offers (Krupka-Weber elicitation) in Ultimatum Game context
Reading fidelity high
Study strength high
n=2658
1.0
Receiving a rejection of an offer from a machine is no less socially appropriate than receiving the same rejection from a human. Decision Quality null_result social appropriateness ratings for receiving rejections
Reading fidelity high
Study strength high
n=2658
1.0
People apply different norms for machines deciding how to split resources but are not opposed to machines enforcing the norms. Decision Quality mixed inferred social norms applied to machine vs. human decision-making and enforcement
Reading fidelity high
Study strength medium
n=2658
0.6
The findings are consistent with offers made by machines being viewed as having both a cognitive and emotional component. Ai Safety And Ethics mixed interpretation regarding perceived components (cognitive and emotional) of machine-made offers
Reading fidelity high
Study strength speculative
n=2658
0.1
The study used the Krupka-Weber norm elicitation task to detect shifts in social appropriateness ratings in Dictator and Ultimatum Games. Research Productivity other method for eliciting social appropriateness norms
Reading fidelity high
Study strength high
n=2658
1.0
Two pre-registered and incentivized experimental studies used representative samples from the UK and US with combined sample size N = 2,658. Research Productivity other study design and sample composition
Reading fidelity high
Study strength high
n=2658
1.0

Notes