1 cumulative citations
View corpus contextPeers financially punish LLM users: in an online experiment, participants destroyed on average 36% of the earnings of workers who relied solely on an LLM, and punishment increased with extent of use; curious asymmetries emerged as denials of use were treated with particular suspicion.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The rapid spread of large language models (LLMs) has raised concerns about the social reactions they provoke. Prior research documents negative attitudes toward AI users, but it remains unclear whether such disapproval translates into costly action. We address this question in a two-phase online experiment (N = 491 Phase II participants; Phase I provided targets) where participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. On average, participants destroyed 36% of the earnings of those who relied exclusively on the model, with punishment increasing monotonically with actual LLM use. Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of "no use" are treated with suspicion. Conversely, at high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Taken together, these findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions.
Summary
Main Finding
People are willing to incur personal costs to punish others for using large language models (LLMs). In a two-phase online experiment (Nphase-II = 491), participants on average destroyed 36% of the earnings of peers who relied exclusively on an LLM. Punishment rises with intensity of actual LLM use, and disclosure interacts with credibility: self-reported non-use is punished more than verified non-use, while at high intensities actual use is punished more than self-reported use.
Key Points
- Behavioral evidence that negative attitudes toward AI users translate into costly antisocial actions (money burning).
- H1 supported: LLM users are punished more than those with no access to LLMs.
- H2 supported: Antisocial behavior increases monotonically with the intensity of actual LLM use.
- H3 supported: Self-reported null use (claiming no LLM use) is punished more harshly than actual, observable non-use — a credibility gap.
- H4 supported (differential slopes): The relationship between punishment and reported use differs from punishment and actual use; at high levels, actual reliance can attract heavier punishment than self-reported reliance.
- Transparency is a double-edged sword: disclosure can backfire when observers distrust the signal.
- Results are preregistered and materials/data/code are publicly available (OSF link in paper).
Data & Methods
- Design: Two-phase experiment.
- Phase I (N = 143 recruited): participants completed five incentivized real-effort emoji-counting tasks under one of three conditions: (1) control — no LLM access; (2) actual LLM access with actual usage observable to Phase II; (3) LLM access with self-reported usage (only the report shown to Phase II). LLM outputs (Google Gemini) were pre-generated and restricted to tasks Gemini solved correctly.
- Phase II (N recruited = 501; final N = 491 after exclusions): representative UK sample from Prolific. Within-subjects: each Phase-II decider made money-burning decisions for 13 target cases (control + actual-use 0–5 + reported-use 0–5). One decision implemented at random.
- Incentives:
- Phase I: £1 flat + up to £1 performance bonus (20p per correct task).
- Phase II: £2.25 flat + 40-token endowment (1 token = £0.01). Each token spent reduced the target’s 100-token bonus by 5 tokens; spending 20 tokens could eliminate the entire bonus. Burning produced no gain to the burner.
- Target selection: only Phase-I participants with perfect task scores were presented to deciders — isolates punishment due to perceived undeservedness rather than performance.
- Measurement: Dependent variable = proportion of target bonus burned by decider (13 choices per decider).
- Statistical approach: Beta-regression mixed models (glmmTMB) on transformed proportional outcome, with random intercepts for participants and block orders and a random slope for intensity. Control variables and robustness checks reported; preregistration noted one minor deviation on attention-check exclusion criteria.
- Open science: preregistration, data/materials/code available on OSF.
Implications for AI Economics
- Social externalities reduce the net private benefits of LLM adoption: observable productivity gains can generate reputational and monetary sanctions by peers, effectively imposing an adoption cost not usually captured in productivity estimates.
- Diffusion dynamics: models of technology adoption should incorporate social-preference-driven friction (punishment, reputational costs) and signaling equilibria. Visibility of use and disclosure rules will materially affect adoption paths.
- Labor markets and incentives:
- Workers may hide LLM use to avoid sanctions, reducing transparency and potentially increasing mismatch between measured and actual productivity.
- Employers should anticipate social frictions when encouraging AI tools; policies (formal disclosure rules, clear norms, team-based incentives, certification of valid AI use) can mitigate stigma and align incentives.
- Welfare trade-offs: efficiency gains from AI might be offset by welfare losses arising from costly norm enforcement. Policymakers and firms should consider institutional mechanisms that reduce false signaling and build credible channels for legitimate AI use (e.g., audit frameworks, verified disclosure, standardized attribution).
- Design of disclosure regimes: Because declarations of non-use can backfire, naive transparency requirements may unintentionally create credibility gaps. Designing credibility-enhancing mechanisms (e.g., verifiable usage logs, attestations, or accepted usage labels) could reduce punishment.
- Research agenda: Extend to higher-stakes, field, and team settings; examine heterogeneity (cultural, sectoral, hierarchical); investigate long-run equilibrium (does punishment reduce AI use or shift it covertly?), and test institutional interventions (verified disclosure, normative campaigns, compensation redesign).
Limitations to bear in mind: lab-style monetary stakes are modest; LLM outputs were static and pre-validated (no real-time errors); Phase-I cell sizes were unequal; UK sample—generalizability across cultures and sectors requires further testing.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We conducted a two-phase online experiment (Phase I provided targets; Phase II had N = 491 participants). Other | null_result | experimental design / sample composition (Phase I targets, Phase II decision-makers) |
Reading fidelity
high
Study strength
high
|
n=491
|
| Participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. Wages | null_result | punishment decisions (amount spent to reduce peers' earnings) |
Reading fidelity
high
Study strength
high
|
n=491
|
| On average, participants destroyed 36% of the earnings of those who relied exclusively on the model. Wages | negative | share of peers' earnings destroyed (punishment amount) |
Reading fidelity
high
Study strength
high
|
n=491
36% of the earnings destroyed
|
| Punishment increased monotonically with actual LLM use. Wages | negative | punishment amount as a function of actual LLM usage level |
Reading fidelity
high
Study strength
high
|
n=491
|
| Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of 'no use' are treated with suspicion. Wages | negative | punishment amount comparing self-reported null use vs actual null use |
Reading fidelity
high
Study strength
high
|
n=491
|
| At high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Wages | negative | punishment amount comparing actual high reliance vs self-reported high reliance |
Reading fidelity
high
Study strength
high
|
n=491
|
| These findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions. Wages | negative | incidence and magnitude of social sanctions (punishment) faced by LLM users |
Reading fidelity
high
Study strength
medium
|
n=491
|