15 cumulative citations
View corpus contextLarge language models match human persuaders on average, but results swing widely by context; a meta-analysis of seven studies finds no average advantage for humans or LLMs, while model choice, message design and domain jointly explain most of the variation.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models (LLMs) are increasingly used for persuasion, such as in political communication and marketing, where they affect how people think, choose, and act. Yet, empirical findings on the effectiveness of LLMs in persuasion compared to humans remain inconsistent. The aim of this study was to systematically review and meta-analytically assess whether LLMs differ from humans in persuasive effectiveness. We identified $7$ studies with 17,422 participants primarily recruited from English-speaking countries and $12$ effect size estimates. Egger's test indicated potential small-study effects ($p = .018$), but the trim-and-fill analysis did not impute any missing studies, suggesting a low risk of publication bias. We then compute the standardized effect sizes based on Hedges' $g$. The results show no significant overall difference in persuasive performance between LLMs and humans ($g = 0.02$, $p = .530$). However, we observe substantial heterogeneity across studies ($I^2 = 75.97\%$), suggesting that persuasiveness strongly depends on contextual factors. In separate exploratory moderator analyses, no individual factor (e.g., LLM model, conversation design, or domain) reached statistical significance, which may be due to the limited number of studies. When considered jointly in a combined model, these factors explained a large proportion of the between-study variance ($R^2 = 81.93\%$), and residual heterogeneity is low ($I^2 = 35.51\%$). Although based on a small number of studies, this suggests that differences in LLM model, conversation design, and domain are important contextual factors in shaping persuasive performance, and that single-factor tests may understate their influence. Our results highlight that LLMs can match human performance in persuasion, but their success depends strongly on how they are implemented and embedded in communication contexts.
Summary
Main Finding
Across 12 effect-size estimates from 7 experimental studies (N = 17,422 participants), large language models (LLMs) overall do not differ significantly from humans in persuasive effectiveness (Hedges’ g = 0.02, p = 0.53). However, there is substantial between-study heterogeneity (I2 ≈ 76%), and contextual factors (LLM model, interaction design, domain) jointly explain a large share of that heterogeneity (combined-model R2 ≈ 81.9%), suggesting that LLM persuasiveness depends strongly on implementation and context.
Key Points
- Evidence base: 7 studies yielding 12 independent comparisons; participants mainly from English-speaking countries (majority recruited via Prolific, mTurk, Lucid).
- Overall effect: No statistically meaningful difference between LLM-generated and human-generated persuasive messages (Hedges’ g = 0.02).
- Heterogeneity: High unexplained heterogeneity in the pooled model (I2 = 75.97%), indicating effect variation across contexts.
- Small-study signals: Egger’s test indicated potential small-study effects (p = 0.018), but trim-and-fill imputed no missing studies (authors interpret this as low publication-bias risk).
- Moderation: Single-factor moderator tests (LLM model, interaction format, domain) were not individually significant (likely low power). A joint model including multiple contextual moderators explained most between-study variance (R2 = 81.93%) and reduced residual heterogeneity (I2 → 35.51%).
- Outcomes & operationalization: Studies used varied outcome measures (attitude change, behavioral intention, compliance, perceived message effectiveness), with the meta-analysis prioritizing behavioral/intention measures when available.
- Limitations: Small number of eligible studies (stringent inclusion criteria), reliance on lab/online experiments, heterogenous outcomes and measures, and most samples from convenience online panels.
Data & Methods
- Scope & search: Systematic review and meta-analysis following PRISMA; searched Web of Science, ACM DL, arXiv, SSRN, OSF; English-language papers after Dec 2022; search cutoff May 22, 2025.
- Inclusion criteria: Direct experimental comparisons between LLM-generated and human-generated messages; between-subjects independent observations at message/participant level; persuasive outcome measure reported; sufficient statistics to compute effect sizes.
- Final corpus: 7 eligible publications; some reported multiple relevant outcomes → 12 effect sizes; total N = 17,422.
- Coding: Extracted metadata and experimental features (LLM type: GPT-3.x, GPT-3.5, GPT-4.x, Claude 3.x; interaction type: one-shot vs interactive; domain: politics, health, other; content length; personalization; recruitment source; outcome type and measurement).
- Effect-size computation: Standardized using Hedges’ g (Cohen’s d converted from available stats and corrected for small-sample bias).
- Meta-analytic model: Random-effects model with REML estimator; heterogeneity quantified via τ2, I2, and Cochran’s Q. Moderator analyses via subgroup tests and meta-regressions.
- Publication-bias checks: Egger’s regression and trim-and-fill; influence diagnostics and leave-one-out sensitivity analyses.
- Reproducibility: Analysis code and data reported as publicly available (per paper).
Implications for AI Economics
- Substitutability and market structure
- On average, LLMs can match human persuaders, implying potential substitution in many persuasion tasks (marketing copywriting, targeted messaging, some policy communication). This raises implications for labor demand in entry-to-mid-level persuasive content roles and for firms’ cost structures in customer acquisition.
- Heterogeneous performance by context means specialization remains valuable—LLMs may displace routine, standardized persuasion work first, while human experts retain advantage in high-context, high-stakes, or highly personalized persuasion.
- Pricing and adoption decisions
- Firms should evaluate cost-effectiveness by context (domain, interactivity, model choice). The joint-moderator result suggests returns to careful implementation (model selection, dialogue format, domain adaptation), so adoption decisions should internalize implementation costs and expected quality gains.
- Platform and competition effects
- Widespread, low-cost LLM persuasion could lower marginal costs of persuasion, increasing supply of persuasive content and potentially altering equilibrium attention markets (greater competition for user attention, higher advertising volumes).
- Entry barriers may fall for new market entrants who can use LLMs to produce persuasive content quickly, affecting incumbents and market concentration dynamics.
- Externalities, misinformation, and regulation
- Because LLMs can be as persuasive as humans in some contexts, scale-up increases risks of manipulation, misinformation, and political influence. Economic models of information markets should incorporate externalities from mass LLM-generated persuasion and consider policy interventions (disclosure requirements, platform liability, differential incentives).
- Welfare and consumer surplus
- Potential welfare gains from personalized, informative persuasion (e.g., health nudges) coexist with harms from deceptive or low-quality persuasion. Evaluations must weigh aggregate welfare changes, distributional impacts (which groups are most targeted), and longer-term preference formation effects.
- Research & evaluation priorities for economists
- Need for larger, preregistered randomized field experiments measuring actual behavioral outcomes (not only attitudes/PME) and cost-effectiveness metrics.
- Estimation of heterogeneous treatment effects (who is most/least susceptible), pricing experiments to assess willingness-to-pay for LLM vs human-produced persuasion, and structural models of labor reallocation and firm adoption dynamics.
- Incorporate model-upgrade dynamics (as LLM capability evolves) into equilibrium analyses: models where persuasion technology improves over time and impacts markets, wages, and regulation.
- Policy design guidance
- Given context sensitivity, regulation might target high-risk domains (political advertising, health misinformation) and require transparency/disclosure for LLM-generated persuasive content, while allowing experimentation in low-risk domains where efficiency gains are clearer.
- Monitoring and auditing (e.g., randomized audits, platform reporting) will be important to quantify real-world impact and guide policy calibration.
Limitations to keep in mind when applying results: the meta-analysis is based on a small, selective set of experiments (7 studies) with heterogeneous measures and mainly online convenience samples; external validity to real-world, large-scale campaigns is still uncertain. Future economic research should prioritize field-level outcome measurement, cost-benefit analyses, and modeling of equilibrium and distributional effects.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The systematic review identified 7 studies with 17,422 participants and 12 effect size estimates. Decision Quality | null_result | number of studies and participants included in the meta-analysis |
Reading fidelity
high
Study strength
high
|
n=17422
7 studies, 17,422 participants, 12 effect size estimates
|
| Egger's test indicated potential small-study effects (p = .018). Decision Quality | mixed | evidence of small-study effects / publication bias |
Reading fidelity
high
Study strength
medium
|
n=7
p = .018
|
| Trim-and-fill analysis did not impute any missing studies, suggesting a low risk of publication bias. Decision Quality | null_result | evidence of publication bias via trim-and-fill |
Reading fidelity
high
Study strength
medium
|
n=7
no missing studies imputed by trim-and-fill
|
| There is no significant overall difference in persuasive performance between LLMs and humans (Hedges' g = 0.02, p = .530). Decision Quality | null_result | persuasive performance (LLMs vs. humans) |
Reading fidelity
high
Study strength
medium
|
n=17422
g = 0.02, p = .530
|
| There is substantial heterogeneity across studies (I^2 = 75.97%), suggesting that persuasiveness strongly depends on contextual factors. Decision Quality | mixed | between-study heterogeneity in effect sizes (I^2) |
Reading fidelity
high
Study strength
high
|
n=7
I^2 = 75.97%
|
| In separate exploratory moderator analyses, no individual factor (e.g., LLM model, conversation design, or domain) reached statistical significance, possibly due to the limited number of studies. Decision Quality | null_result | moderator effects of LLM model, conversation design, and domain on persuasive effectiveness |
Reading fidelity
high
Study strength
low
|
n=7
no individual moderator reached statistical significance (no p-values reported here)
|
| When considered jointly in a combined model, LLM model, conversation design, and domain explained a large proportion of the between-study variance (R^2 = 81.93%), and residual heterogeneity was low (I^2 = 35.51%). Decision Quality | positive | proportion of between-study variance explained by combined moderators (R^2) and residual heterogeneity (I^2) |
Reading fidelity
high
Study strength
low
|
n=7
R^2 = 81.93%, residual I^2 = 35.51%
|
| Although based on a small number of studies, the results suggest that differences in LLM model, conversation design, and domain are important contextual factors in shaping persuasive performance, and that single-factor tests may understate their influence. Decision Quality | mixed | importance of implementation/contextual factors (LLM model, conversation design, domain) for persuasive performance |
Reading fidelity
high
Study strength
low
|
n=7
interpretive statement based on meta-analytic heterogeneity and combined-moderator R^2
|
| The results highlight that LLMs can match human performance in persuasion, but their success depends strongly on how they are implemented and embedded in communication contexts. Decision Quality | null_result | comparative persuasive effectiveness of LLMs versus humans |
Reading fidelity
high
Study strength
medium
|
n=17422
g = 0.02 (non-significant); contextual dependence inferred from heterogeneity/moderators
|
| Participants in the included studies were primarily recruited from English-speaking countries. Decision Quality | null_result | geographic/language provenance of participant samples |
Reading fidelity
high
Study strength
medium
|
n=17422
primarily recruited from English-speaking countries
|