The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-written ads beat human copy in lab preference tests, winning 59% to 41% overall and dominating authority and consensus appeals; when ads are tailored to personality, AI performs on par with human experts.

LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
Elyas Meguellati, Stefano Civelli, Lei Han, Abraham Bernstein, Shazia Sadiq, Gianluca Demartini · December 03, 2025
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Elyas Meguellati unresolved corpus identity
  2. Stefano Civelli unresolved corpus identity
  3. Lei Han unresolved corpus identity
  4. Abraham Bernstein unresolved corpus identity
  5. Shazia Sadiq unresolved corpus identity
  6. Gianluca Demartini unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Elyas Meguellati provider ID
  2. Stefano Civelli provider ID
  3. Lei Han provider ID
  4. Abraham Bernstein provider ID
  5. S. Sadiq provider ID
  6. Gianluca Demartini provider ID
In controlled experiments, LLM-generated ads matched humans on personality-targeted messaging and outperformed human-created ads on universal persuasion frames—winning 59.1% of preference choices overall, especially on authority and consensus appeals.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As large language models (LLMs) become increasingly capable of generating persuasive content, understanding their effectiveness across different advertising strategies becomes critical. This paper presents a two-part investigation examining LLM-generated advertising through complementary lenses: (1) personality-based and (2) psychological persuasion principles. In our first study (n=400), we tested whether LLMs could generate personalized advertisements tailored to specific personality traits (openness and neuroticism) and how their performance compared to human experts. Results showed that LLM-generated ads achieved statistical parity with human-written ads (51.1% vs. 48.9%, p > 0.05), with no significant performance differences for matched personalities. Building on these insights, our second study (n=800) shifted focus from individual personalization to universal persuasion, testing LLM performance across four foundational psychological principles: authority, consensus, cognition, and scarcity. AI-generated ads significantly outperformed human-created content, achieving a 59.1% preference rate (vs. 40.9%, p < 0.001), with the strongest performance in authority (63.0%) and consensus (62.5%) appeals. Qualitative analysis revealed AI's advantage stems from crafting more sophisticated, aspirational messages and achieving superior visual-narrative coherence. Critically, this quality advantage proved robust: even after applying a 21.2 percentage point detection penalty when participants correctly identified AI-origin, AI ads still outperformed human ads, and 29.4% of participants chose AI content despite knowing its origin. These findings demonstrate LLMs' evolution from parity in personalization to superiority in persuasive storytelling, with significant implications for advertising practice given LLMs' near-zero marginal cost and time requirements compared to human experts.

Summary

Main Finding

LLM-generated ads match humans on trait-targeted personalization but outperform humans when deploying universal persuasion principles. In two experiments (Study 1: n=400; Study 2: n=800) the authors find parity for personality-tailored ads (LLM preference 51.1% vs human 48.9%, p>0.05) but clear AI superiority for persuasion-based ads (LLM 59.1% vs human 40.9%, p<0.001). The AI advantage is strongest for authority (63.0%) and consensus (62.5%) appeals and remains robust after accounting for participants’ detection of AI origin.

Key Points

  • Two-part empirical design:
    • Study 1: personality-based personalization (openness, neuroticism). Measures: Likert ratings (product attitude, purchase intention, engagement intention) and side-by-side choice between LLM vs human ads.
    • Study 2: universal persuasion principles (authority, consensus, cognition/processing fluency, scarcity) with multimodal ads and qualitative follow-up.
  • Outcomes:
    • Personalization parity: no significant difference between LLM- and human-generated ads for matched personality targets (51.1% vs 48.9%, p > 0.05).
    • Persuasion superiority: LLM ads preferred substantially more than human ads (59.1% vs 40.9%, p < 0.001); largest gains in authority and consensus strategies.
  • Qualitative diagnostics: participants cited more sophisticated, aspirational messaging and better visual–narrative coherence in AI ads as drivers of preference.
  • Origin effects: even when adjusting preferences downward by a 21.2 percentage-point “detection penalty” for correctly identified AI ads, AI content still outperformed human content; 29.4% of participants selected AI content despite knowing it was AI-generated.
  • Generation pipeline: authors combine Knowledge Graph grounding (mapping user traits to product attributes) with LLM “generative storytelling” to produce context-driven ad copy; experiments isolate the generative module (no live deployment confounds).
  • Practical advantage: LLMs offer near-zero marginal cost and low latency compared with expert human creation, implying potential large-scale production and rapid iteration.

Data & Methods

  • Experimental design:
    • Study 1 (n=400): participants exposed to ads targeted for high-openness or high-neuroticism profiles. Tasks: (1) rate ad on 3 Likert outcomes (product attitude, purchase intention, engagement intention); (2) choose preferred ad in side-by-side human vs LLM comparison. Personality measured by a 20-item Big Five questionnaire administered after ad tasks.
    • Study 2 (n=800): participants evaluated LLM- vs human-crafted ads engineered around four persuasion principles (authority, consensus, cognition/fluency, scarcity). Included multimodal stimuli and qualitative prompts capturing reasons for preference and recognition of content origin.
  • Generation approach:
    • Knowledge Graphs used to ground prompts (map trait → product attributes), then LLM(s) used to generate ad copy and visual–narrative elements. Human ads drawn from expert-created materials (sourced from existing creative/agency outputs—paper provides sourcing details).
  • Outcome metrics and analysis:
    • Primary metric: binary preference (which ad chosen). Secondary metrics: Likert ratings for affect and behavioral intent, qualitative coding of reasons.
    • Statistical tests: preference share comparisons with p-values reported (e.g., p<.001 for persuasion superiority). Robustness check applying a detection-penalty adjustment (21.2 percentage points) to account for origin recognition bias.
  • Qualitative methods: thematic analysis identifying motifs (aspirational tone, coherence) explaining AI advantage.

Implications for AI Economics

  • Cost structure and scale:
    • Near-zero marginal cost and faster turnaround for LLM-generated creative imply dramatically lower marginal production costs for advertising assets, enabling massive scaling of ad variants (combinatorial personalization and A/B testing).
  • Labor & task allocation:
    • Creative production jobs (copywriters, junior creatives) face displacement or role-shift; human labor may concentrate on high-level strategy, brand stewardship, legal/ethical oversight, and curation/editing rather than first-draft generation.
    • A premium may persist for rare human skills (brand voice, deep domain expertise), but routine creative iteration becomes automatable.
  • Industry structure & competition:
    • Lower barriers to producing persuasive creative could intensify competition among advertisers (more variants, faster optimization), shifting industry dynamics toward firms that can best integrate LLMs into pipelines and measure ROI.
    • Advertising agencies’ value proposition could pivot to strategy, sourcing, and governance rather than execution; platform and tool providers that bundle KG-grounding, model access, and creative workflows may capture rents.
  • Pricing, auctions, and ad markets:
    • If LLM ads yield higher click/engagement rates, this will affect auction dynamics (higher bids justified by higher expected returns), potentially raising equilibrium prices for premium placements unless supply of effective creatives expands even faster.
    • Increased ad effectiveness could raise advertisers’ willingness to pay per impression/engagement, altering allocation of ad spend across channels.
  • Consumer welfare, externalities, and regulation:
    • Higher persuasive potency (especially using authority/consensus cues) raises concerns about manipulation, misinformation, and consumer protection—calling for updated disclosure, transparency rules, and possibly limits on certain persuasion techniques.
    • Privacy and microtargeting ethics remain central: cheap personalized persuasion amplifies risks of exploitative targeting, nudging, and political-messaging externalities.
  • Measurement & market power:
    • Firms that own large user data, KGs, and fine-tuned creative pipelines can extract more value from LLMs, enhancing platform advantages and potential concentration.
  • Research and policy needs:
    • Need for field experiments measuring real behavioral outcomes (purchases, long-run engagement), cross-platform effects, and welfare implications.
    • Policy considerations: mandatory disclosures, limits on manipulative persuasion in sensitive domains, auditability of persuasion pipelines, and support for workforce transition.
  • Investment and ROI:
    • Expect reallocation of media budgets: more spending on dynamic creative optimization and measurement; less on high-cost per-unit creative production. Returns to efficient experimentation and measurement tools increase.

Limitations & open questions (relevant for economic analysis) - External validity: lab/online preference measures do not directly translate to long-run purchase behavior or market-level outcomes. - Product & context scope: effectiveness may vary by product category, regulatory environment, and cultural context. - Model & pipeline heterogeneity: results depend on LLM quality, prompt engineering, and KG grounding—returns will vary across implementations. - Origin-disclosure dynamics: real-world effects of disclosure, reputational impacts, and changing consumer priors require longitudinal study.

If you want, I can: - Convert these implications into a short memo targeted to agency executives or policymakers. - Produce a simple economic model (graphical or algebraic) of how marginal cost reductions in creative affect ad-auction equilibrium and creative labor demand.

Assessment

Paper Typerct Evidence Strengthmedium — The randomized experiments give credible causal identification for effects on short-term stated preferences (internal validity is good), but the outcome is preference in an experiment rather than real-world behavioral or economic outcomes (clicks, conversions, sales), and external validity is limited by sample and ecological constraints. Methods Rigormedium — Studies have reasonably large sample sizes (n=400 and n=800), direct comparison with human experts, statistical testing, and robustness checks (detection penalty, qualitative coding). However, the description lacks detail on sampling frame and recruitment, pre-registration, blinding, the representativeness and selection of human copywriters, the exact randomization protocol, and downstream behavioral/market validation. SampleTwo online experimental samples: Study 1 (n=400) comparing LLM vs human ads targeted to personality traits (openness, neuroticism); Study 2 (n=800) comparing LLM vs human ads across four persuasion principles (authority, consensus, cognition, scarcity); participants were asked preference judgments and whether they identified ad origin; ad content produced by a specified LLM and by human experts (details on platform/participant recruitment, demographics, product contexts, and LLM version not reported here). Themesproductivity human_ai_collab IdentificationRandomized controlled online experiments: participants were randomly shown AI-generated versus human-created advertisements and asked to indicate preferences; Study 1 tested ads targeted to measured personality traits (openness, neuroticism) and compared matched performance, Study 2 randomized persuasion-frame (authority, consensus, cognition, scarcity) and compared preference shares; robustness checks included a detection-penalty adjustment and qualitative coding of message features. GeneralizabilityStated preference in an online experiment may not translate to real-world consumer behavior (clicks, purchases, long-term engagement), Likely convenience/online sample (e.g., MTurk/Prolific) so not nationally representative, Results may depend on the specific LLM model, prompt engineering, and human copywriter pool used, Limited range of product categories, brands, channels, and cultural contexts reduces external validity, Short-term single-exposure tests do not capture campaign dynamics, ad fatigue, or downstream economic effects

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In Study 1 (n=400), LLM-generated ads achieved statistical parity with human-written ads (51.1% vs. 48.9%, p > 0.05). Consumer Welfare null_result ad preference rate
Reading fidelity high
Study strength high
n=400
51.1% vs. 48.9%, p > 0.05
1.0
In Study 1 there were no significant performance differences for ads matched to participants' personalities (openness and neuroticism). Consumer Welfare null_result ad preference rate by personality match
Reading fidelity medium
Study strength medium
not reported
0.36
In Study 2 (n=800), AI-generated ads significantly outperformed human-created content, achieving a 59.1% preference rate (vs. 40.9%, p < 0.001). Consumer Welfare positive ad preference rate
Reading fidelity high
Study strength high
n=800
59.1% vs. 40.9%, p < 0.001
1.0
AI-generated ads showed their strongest performance in authority-based persuasion (63.0% preference). Consumer Welfare positive ad preference rate for authority appeals
Reading fidelity high
Study strength medium
63.0% (authority appeals)
0.6
AI-generated ads also strongly outperformed on consensus-based persuasion (62.5% preference). Consumer Welfare positive ad preference rate for consensus appeals
Reading fidelity high
Study strength medium
62.5% (consensus appeals)
0.6
Qualitative analysis indicated AI's advantage stems from crafting more sophisticated, aspirational messages and achieving superior visual–narrative coherence. Output Quality positive message sophistication and visual-narrative coherence (qualitative)
Reading fidelity medium
Study strength medium
not reported
0.36
Even after applying a 21.2 percentage point detection penalty when participants correctly identified AI origin, AI ads still outperformed human ads; 29.4% of participants chose AI content despite knowing its origin. Consumer Welfare positive ad preference rate after detection/attribution adjustment
Reading fidelity high
Study strength high
21.2 percentage point detection penalty; 29.4% chose AI despite knowing origin
1.0
LLMs have near-zero marginal cost and time requirements compared to human experts, which has significant implications for advertising practice. Organizational Efficiency positive marginal cost and time requirements
Reading fidelity medium
Study strength speculative
not reported
0.06
Overall conclusion: LLMs have evolved from parity in personality-based personalization (Study 1) to superiority in persuasive storytelling/universal persuasion (Study 2). Consumer Welfare mixed relative effectiveness (personalization vs. persuasive storytelling)
Reading fidelity high
Study strength medium
n=1200
0.6

Notes