The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Peers financially punish LLM users: in an online experiment, participants destroyed on average 36% of the earnings of workers who relied solely on an LLM, and punishment increased with extent of use; curious asymmetries emerged as denials of use were treated with particular suspicion.

Antisocial behavior towards large language model users: experimental evidence
Paweł Niszczota, Cassandra Grützner · January 14, 2026
arxiv rct high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Paweł Niszczota unresolved corpus identity
  2. Cassandra Grützner unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Pawel Niszczota provider ID
  2. Cassandra Grützner provider ID
In an incentivized online experiment, peers punished workers who used LLMs—on average destroying 36% of earnings for exclusive LLM users—with punishment rising with degree of actual use and showing systematic differences between self-reported and actual use.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The rapid spread of large language models (LLMs) has raised concerns about the social reactions they provoke. Prior research documents negative attitudes toward AI users, but it remains unclear whether such disapproval translates into costly action. We address this question in a two-phase online experiment (N = 491 Phase II participants; Phase I provided targets) where participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. On average, participants destroyed 36% of the earnings of those who relied exclusively on the model, with punishment increasing monotonically with actual LLM use. Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of "no use" are treated with suspicion. Conversely, at high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Taken together, these findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions.

Summary

Main Finding

People are willing to incur personal costs to punish others for using large language models (LLMs). In a two-phase online experiment (Nphase-II = 491), participants on average destroyed 36% of the earnings of peers who relied exclusively on an LLM. Punishment rises with intensity of actual LLM use, and disclosure interacts with credibility: self-reported non-use is punished more than verified non-use, while at high intensities actual use is punished more than self-reported use.

Key Points

  • Behavioral evidence that negative attitudes toward AI users translate into costly antisocial actions (money burning).
  • H1 supported: LLM users are punished more than those with no access to LLMs.
  • H2 supported: Antisocial behavior increases monotonically with the intensity of actual LLM use.
  • H3 supported: Self-reported null use (claiming no LLM use) is punished more harshly than actual, observable non-use — a credibility gap.
  • H4 supported (differential slopes): The relationship between punishment and reported use differs from punishment and actual use; at high levels, actual reliance can attract heavier punishment than self-reported reliance.
  • Transparency is a double-edged sword: disclosure can backfire when observers distrust the signal.
  • Results are preregistered and materials/data/code are publicly available (OSF link in paper).

Data & Methods

  • Design: Two-phase experiment.
    • Phase I (N = 143 recruited): participants completed five incentivized real-effort emoji-counting tasks under one of three conditions: (1) control — no LLM access; (2) actual LLM access with actual usage observable to Phase II; (3) LLM access with self-reported usage (only the report shown to Phase II). LLM outputs (Google Gemini) were pre-generated and restricted to tasks Gemini solved correctly.
    • Phase II (N recruited = 501; final N = 491 after exclusions): representative UK sample from Prolific. Within-subjects: each Phase-II decider made money-burning decisions for 13 target cases (control + actual-use 0–5 + reported-use 0–5). One decision implemented at random.
  • Incentives:
    • Phase I: £1 flat + up to £1 performance bonus (20p per correct task).
    • Phase II: £2.25 flat + 40-token endowment (1 token = £0.01). Each token spent reduced the target’s 100-token bonus by 5 tokens; spending 20 tokens could eliminate the entire bonus. Burning produced no gain to the burner.
  • Target selection: only Phase-I participants with perfect task scores were presented to deciders — isolates punishment due to perceived undeservedness rather than performance.
  • Measurement: Dependent variable = proportion of target bonus burned by decider (13 choices per decider).
  • Statistical approach: Beta-regression mixed models (glmmTMB) on transformed proportional outcome, with random intercepts for participants and block orders and a random slope for intensity. Control variables and robustness checks reported; preregistration noted one minor deviation on attention-check exclusion criteria.
  • Open science: preregistration, data/materials/code available on OSF.

Implications for AI Economics

  • Social externalities reduce the net private benefits of LLM adoption: observable productivity gains can generate reputational and monetary sanctions by peers, effectively imposing an adoption cost not usually captured in productivity estimates.
  • Diffusion dynamics: models of technology adoption should incorporate social-preference-driven friction (punishment, reputational costs) and signaling equilibria. Visibility of use and disclosure rules will materially affect adoption paths.
  • Labor markets and incentives:
    • Workers may hide LLM use to avoid sanctions, reducing transparency and potentially increasing mismatch between measured and actual productivity.
    • Employers should anticipate social frictions when encouraging AI tools; policies (formal disclosure rules, clear norms, team-based incentives, certification of valid AI use) can mitigate stigma and align incentives.
  • Welfare trade-offs: efficiency gains from AI might be offset by welfare losses arising from costly norm enforcement. Policymakers and firms should consider institutional mechanisms that reduce false signaling and build credible channels for legitimate AI use (e.g., audit frameworks, verified disclosure, standardized attribution).
  • Design of disclosure regimes: Because declarations of non-use can backfire, naive transparency requirements may unintentionally create credibility gaps. Designing credibility-enhancing mechanisms (e.g., verifiable usage logs, attestations, or accepted usage labels) could reduce punishment.
  • Research agenda: Extend to higher-stakes, field, and team settings; examine heterogeneity (cultural, sectoral, hierarchical); investigate long-run equilibrium (does punishment reduce AI use or shift it covertly?), and test institutional interventions (verified disclosure, normative campaigns, compensation redesign).

Limitations to bear in mind: lab-style monetary stakes are modest; LLM outputs were static and pre-validated (no real-time errors); Phase-I cell sizes were unequal; UK sample—generalizability across cultures and sectors requires further testing.

Assessment

Paper Typerct Evidence Strengthhigh — Causal effects are identified via an experimental design with monetary incentives and direct behavioral measures of punishment; sample size for the punishment stage (N=491) is reasonably large for an online experiment and the treatment variation (actual vs. reported LLM use) maps cleanly onto the outcomes. Methods Rigormedium — Design shows strong internal validity (randomized treatments, real monetary stakes), but potential limitations include online convenience sampling, limited ecological validity of a one-off lablike task, possible demand effects or signaling artifacts around disclosure, and (from the summary) no detail on pre-registration, manipulation checks, heterogeneity checks, or robustness to alternative specifications. SampleTwo-phase online sample: Phase I produced targets who completed a real-effort task with varying levels of LLM assistance (actual use and self-reported disclosure); Phase II comprised 491 online participants who received an endowment and could pay to reduce targets' earnings; participants were recruited online (platform unspecified) and decisions involved real monetary incentives. Themeshuman_ai_collab labor_markets IdentificationTwo-phase online randomized experiment: Phase I produced peers (targets) who completed a real-effort task with varying, experimentally-manipulated levels of LLM support and/or self-reported disclosure; Phase II participants were given an incentivized endowment and randomly assigned to observe targets' reported/actual LLM use and could spend their endowment to reduce targets' earnings, allowing causal estimation of the effect of observed/actual LLM use on costly punishment. GeneralizabilityOnline convenience sample may not represent broader worker populations or employers, Single-task real-effort setting may not capture complex, collaborative, or high-stakes workplace tasks, Short-term, anonymous experimental interactions may over- or under-state sanctions compared with persistent workplace relationships, Cultural or platform-specific norms (unspecified) may limit cross-country generalizability, Magnitude of monetary stakes in the experiment may differ from real-world costs and incentives, Findings relate to observed/presented use of one class of LLMs and may not generalize to other AI tools or evolving disclosure norms

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We conducted a two-phase online experiment (Phase I provided targets; Phase II had N = 491 participants). Other null_result experimental design / sample composition (Phase I targets, Phase II decision-makers)
Reading fidelity high
Study strength high
n=491
1.0
Participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. Wages null_result punishment decisions (amount spent to reduce peers' earnings)
Reading fidelity high
Study strength high
n=491
1.0
On average, participants destroyed 36% of the earnings of those who relied exclusively on the model. Wages negative share of peers' earnings destroyed (punishment amount)
Reading fidelity high
Study strength high
n=491
36% of the earnings destroyed
1.0
Punishment increased monotonically with actual LLM use. Wages negative punishment amount as a function of actual LLM usage level
Reading fidelity high
Study strength high
n=491
1.0
Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of 'no use' are treated with suspicion. Wages negative punishment amount comparing self-reported null use vs actual null use
Reading fidelity high
Study strength high
n=491
1.0
At high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Wages negative punishment amount comparing actual high reliance vs self-reported high reliance
Reading fidelity high
Study strength high
n=491
1.0
These findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions. Wages negative incidence and magnitude of social sanctions (punishment) faced by LLM users
Reading fidelity high
Study strength medium
n=491
0.6

Notes