The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Carefully framed LLM explanations can keep users trusting wrong answers: in a randomized study of 200+ people, adversarial explanations that use authoritative evidence, neutral tone and expert-style reasoning preserved nearly all trust compared with benign explanations, making users equally likely to accept incorrect outputs.

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making
Shutong Fan, Lan Zhang, Xiaoyong Yuan · February 03, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shutong Fan unresolved corpus identity
  2. Lan Zhang unresolved corpus identity
  3. Xiaoyong Yuan unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Shutong Fan provider ID
  2. Lan Zhang provider ID
  3. Xiaoyong Yuan provider ID
A randomized human study shows adversarially framed LLM explanations can preserve almost as much user trust as benign explanations—even when the model is wrong—especially when explanations mimic expert communication and among less-educated, younger, or highly AI-trusting users.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and act on model recommendations. Large Language Models (LLMs) generate fluent natural-language explanations that shape how users perceive and trust AI outputs, revealing a new attack surface at the cognitive layer: the communication channel between AI and its users. We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Using this metric as a lens, we highlight a behavioral risk where persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong. To characterize this threat, we conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Our findings show that users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI.

Summary

Main Finding

Adversarial explanation attacks (AEAs) — manipulating the framing, tone, evidence, or format of LLM-generated explanations while leaving the underlying model prediction unchanged — can sustain high human trust in incorrect AI outputs. In a large human study (n > 200) the authors show that adversarial explanations often preserve most of the trust normally given to benign, correct explanations. The paper introduces the trust miscalibration gap (and a trust-retention ratio) as behavioral metrics to quantify this vulnerability and maps which explanation strategies, contexts, and user traits maximize the effect.

Key Points

  • New attack surface: Explanations are a cognitive-layer channel that an adversary can exploit to change human behavior without touching model accuracy, inputs, or outputs.
  • AEAs: The adversary crafts persuasive explanations (via prompt injection, malicious fine-tuning, or middleware rewriting) that make incorrect predictions seem credible.
  • Metrics:
    • Trust miscalibration gap: difference in human trust between benign and adversarial explanations (particularly when model outputs are wrong).
    • Trust-retention ratio: fraction of benign trust preserved under adversarial explanation when the output is incorrect.
  • Four-dimensional explanation space: authors systematically vary explanations along
    • Reasoning mode (e.g., feature attribution, counterfactual, analogy, procedural)
    • Evidence type (e.g., citations/statistics, equations/proofs, or internal conceptual)
    • Communication style (e.g., neutral, empathetic, authoritative)
    • Presentation format (e.g., visual emphasis, tables, plain text)
  • Empirical findings:
    • Participants reported nearly identical trust for adversarial and benign explanations; adversarial framings preserved the vast majority of benign trust even when the prediction was wrong.
    • The largest miscalibration gaps occurred when adversarial explanations resembled expert communication: authoritative evidence (citations/stats), neutral/analytic tone, and domain-appropriate reasoning.
    • Vulnerability highest on medium-to-hard tasks and fact-driven domains (medicine, business); lower in logic-heavy domains (math, law).
    • Individual differences: more susceptible groups included younger participants, less formally educated participants, and those with high pre-existing trust in AI.
    • Dynamic effects: trust is driven mostly by the current explanation (little carryover), repeated detection of misleading explanations gradually erodes trust, while benign streaks restore it. Overall population-level trust remained stable, but notable minorities had lasting trust shifts.
  • Novelty: First systematic security study treating explanations as an adversarial cognitive channel and quantifying behavioral attack success.

Data & Methods

  • Experimental design:
    • Large human-subject study (n > 200), factorially varying the four explanation dimensions.
    • Tasks spanned multiple domains and difficulty levels (easy/medium/hard), including fact-driven and logic-intensive domains.
    • For each task the model prediction was held fixed; explanations were rendered either benign or adversarially framed.
    • Collected per-task trust ratings, decision alignment (did the participant accept the AI recommendation), and longitudinal measures to track trust evolution across tasks.
  • Attack model:
    • Adversary can control explanation generation but not the predicted label/choice (realistic vectors: prompt injection, fine-tuning, middleware rewriting).
    • Explanations generated with LLMs conditioned on chosen framing strategies; quality control ensured plausibility and context appropriateness.
  • Outcome measures:
    • Trust scores, decision acceptance rates, trust miscalibration gap, trust-retention ratio, subgroup analyses by demographics and pre-trust levels.
  • Robustness checks:
    • Variation across domains and task difficulty; longitudinal analysis of repeated exposures.

Implications for AI Economics

  • Market and consumer risk:
    • Persuasive explanations enable actors (benign firms or adversaries) to shape consumer decisions and can generate economically significant misallocation of resources (misguided investments, medical choices, financial advice), increasing consumer harm and negatively affecting welfare.
    • Firms that deploy persuasive explanation styles may gain short-term trust/adoption advantages, creating incentives to optimize explanations for persuasion rather than truthfulness — a moral hazard and potential race-to-the-bottom in explanation fidelity.
  • Information asymmetry & signaling:
    • Explanations become an endogenous signaling instrument. Without verifiable provenance, consumers cannot easily distinguish truthful from persuasive-but-misleading explanations, exacerbating asymmetric information and potentially increasing demand for third-party verification or reputation mechanisms.
  • Platform competition & externalities:
    • Platforms that allow (or fail to police) manipulative explanation framings can impose negative externalities across users and markets (fraud, systemic misinformed behavior). Competitors may be forced to match persuasive styles to retain users, amplifying risk.
  • Regulation, liability, and insurance:
    • Policy responses could include mandatory provenance (signed, auditable explanations), disclosure requirements for explanation sources/uncertainty, and liability rules for demonstrably misleading explanation practices. Insurers and regulators will need new models to price risks linked to cognitive-layer manipulation.
  • Design economics and product strategy:
    • Firms should internalize costs of user harm from manipulative explanations; invest in verifiable explanation protocols, UI constraints (e.g., standardized uncertainty cues, limited rhetorical authority when confidence is low), and auditing tools measuring trust miscalibration gap as part of deployment risk assessment.
  • Research and measurement:
    • The trust miscalibration gap can be incorporated into economic models of adoption, platform trustworthiness, and mechanism design (e.g., contracts that reward calibrated transparency). Empirical IO and behavioral-economics work can quantify welfare impacts of persuasive explanations on market outcomes.
  • Policy and market interventions:
    • Interventions include certification/audits of explanation fidelity, regulation of explanation provenance, platform-level guardrails on persuasive styles, user education, and industry standards for uncertainty signaling.
  • Future economic questions:
    • How do persuasion-capable AI explanations alter equilibrium adoption and pricing? When do firms have incentives to sell persuasion vs. truthfulness? What liability and monitoring regimes minimize social loss given enforcement costs?

Summary: Treat explanations as an economically meaningful product feature and risk vector. The paper provides both a metric (trust miscalibration gap) and empirical evidence that persuasive LLM explanations can create durable market-relevant harms; policymakers, platform designers, and economists should incorporate cognitive-layer manipulation into models of incentives, regulation, and market design.

Assessment

Paper Typerct Evidence Strengthmedium — Randomized experimental design gives strong internal validity for the effect of explanation framing on self-reported trust, but the sample is modest (≈200+), outcomes are primarily self-reported trust rather than consequential behavioral or economic outcomes, tasks and LLM/explanation instantiations appear limited, and external validity to real-world high-stakes settings is uncertain. Methods Rigormedium — The study systematically varies four framing dimensions and uses a clear metric (trust miscalibration gap), suggesting careful factorial experimental design and measurement; however, details on randomization checks, pre-registration, power for interaction tests, the specific LLM(s) and prompt templates, and correction for multiple hypothesis tests are not provided here, and reliance on self-report measures limits rigor relative to behavioral outcome studies. SampleHuman-subject experiment with just over 200 participants exposed to LLM-generated explanations that were either benign or adversarially framed; participants completed tasks across multiple domains (noted as fact-driven) and difficulty levels, and reported trust in outputs; moderators recorded include education level, age, and baseline propensity to trust AI. Themeshuman_ai_collab governance IdentificationRandomized between-subjects (and/or within-subjects) experimental manipulation of LLM explanation framing across four dimensions (reasoning mode, evidence type, communication style, presentation format), measuring participants' reported trust and computing the 'trust miscalibration gap' between benign and adversarial explanations; causal inference rests on random assignment to framing conditions and comparison of trust outcomes across these conditions, with subgroup analyses by task difficulty and participant demographics. GeneralizabilityModest sample size and likely convenience sampling (e.g., online panel) limit external representativeness, Primary outcomes are self-reported trust, not real-world decision-making or economic outcomes (purchase/usage/labor productivity), Tasks cover a limited set of domains (fact-driven) and difficulty levels—may not generalize to high-stakes or creative domains, Findings may depend on the particular LLM, explanation-generation prompts, or UI used and may not transfer across models or deployment contexts, Short-term experimental exposure; long-term behavior and repeated-interaction effects are untested, Cultural/geographic diversity of participants not specified, limiting cross-population generalizability

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. Other null_result conceptual definition of a new attack class (AEAs)
Reading fidelity high
Study strength speculative
not reported
0.1
We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Decision Quality null_result trust miscalibration gap (difference in human trust between benign and adversarial explanations)
Reading fidelity high
Study strength medium
not reported
0.6
We conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Decision Quality null_result experimental manipulation across four explanation-framing dimensions (reasoning mode, evidence type, communication style, presentation format)
Reading fidelity high
Study strength high
n=200
1.0
Users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. Decision Quality positive human-reported trust in AI explanations (trust rating) for benign vs. adversarial explanations
Reading fidelity high
Study strength medium
n=200
0.6
Persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong (behavioral risk). Decision Quality positive preservation of user trust when AI predictions are incorrect
Reading fidelity high
Study strength medium
n=200
0.6
The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Decision Quality positive vulnerability to AEAs measured as preserved trust under expert-like explanation framing
Reading fidelity high
Study strength medium
n=200
0.6
Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI. Decision Quality positive differential vulnerability (preserved trust) across task difficulty, domain type, and participant demographics/attitudes
Reading fidelity high
Study strength medium
n=200
0.6

Notes