3 cumulative citations
View corpus contextCarefully framed LLM explanations can keep users trusting wrong answers: in a randomized study of 200+ people, adversarial explanations that use authoritative evidence, neutral tone and expert-style reasoning preserved nearly all trust compared with benign explanations, making users equally likely to accept incorrect outputs.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and act on model recommendations. Large Language Models (LLMs) generate fluent natural-language explanations that shape how users perceive and trust AI outputs, revealing a new attack surface at the cognitive layer: the communication channel between AI and its users. We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Using this metric as a lens, we highlight a behavioral risk where persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong. To characterize this threat, we conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Our findings show that users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI.
Summary
Main Finding
Adversarial explanation attacks (AEAs) — manipulating the framing, tone, evidence, or format of LLM-generated explanations while leaving the underlying model prediction unchanged — can sustain high human trust in incorrect AI outputs. In a large human study (n > 200) the authors show that adversarial explanations often preserve most of the trust normally given to benign, correct explanations. The paper introduces the trust miscalibration gap (and a trust-retention ratio) as behavioral metrics to quantify this vulnerability and maps which explanation strategies, contexts, and user traits maximize the effect.
Key Points
- New attack surface: Explanations are a cognitive-layer channel that an adversary can exploit to change human behavior without touching model accuracy, inputs, or outputs.
- AEAs: The adversary crafts persuasive explanations (via prompt injection, malicious fine-tuning, or middleware rewriting) that make incorrect predictions seem credible.
- Metrics:
- Trust miscalibration gap: difference in human trust between benign and adversarial explanations (particularly when model outputs are wrong).
- Trust-retention ratio: fraction of benign trust preserved under adversarial explanation when the output is incorrect.
- Four-dimensional explanation space: authors systematically vary explanations along
- Reasoning mode (e.g., feature attribution, counterfactual, analogy, procedural)
- Evidence type (e.g., citations/statistics, equations/proofs, or internal conceptual)
- Communication style (e.g., neutral, empathetic, authoritative)
- Presentation format (e.g., visual emphasis, tables, plain text)
- Empirical findings:
- Participants reported nearly identical trust for adversarial and benign explanations; adversarial framings preserved the vast majority of benign trust even when the prediction was wrong.
- The largest miscalibration gaps occurred when adversarial explanations resembled expert communication: authoritative evidence (citations/stats), neutral/analytic tone, and domain-appropriate reasoning.
- Vulnerability highest on medium-to-hard tasks and fact-driven domains (medicine, business); lower in logic-heavy domains (math, law).
- Individual differences: more susceptible groups included younger participants, less formally educated participants, and those with high pre-existing trust in AI.
- Dynamic effects: trust is driven mostly by the current explanation (little carryover), repeated detection of misleading explanations gradually erodes trust, while benign streaks restore it. Overall population-level trust remained stable, but notable minorities had lasting trust shifts.
- Novelty: First systematic security study treating explanations as an adversarial cognitive channel and quantifying behavioral attack success.
Data & Methods
- Experimental design:
- Large human-subject study (n > 200), factorially varying the four explanation dimensions.
- Tasks spanned multiple domains and difficulty levels (easy/medium/hard), including fact-driven and logic-intensive domains.
- For each task the model prediction was held fixed; explanations were rendered either benign or adversarially framed.
- Collected per-task trust ratings, decision alignment (did the participant accept the AI recommendation), and longitudinal measures to track trust evolution across tasks.
- Attack model:
- Adversary can control explanation generation but not the predicted label/choice (realistic vectors: prompt injection, fine-tuning, middleware rewriting).
- Explanations generated with LLMs conditioned on chosen framing strategies; quality control ensured plausibility and context appropriateness.
- Outcome measures:
- Trust scores, decision acceptance rates, trust miscalibration gap, trust-retention ratio, subgroup analyses by demographics and pre-trust levels.
- Robustness checks:
- Variation across domains and task difficulty; longitudinal analysis of repeated exposures.
Implications for AI Economics
- Market and consumer risk:
- Persuasive explanations enable actors (benign firms or adversaries) to shape consumer decisions and can generate economically significant misallocation of resources (misguided investments, medical choices, financial advice), increasing consumer harm and negatively affecting welfare.
- Firms that deploy persuasive explanation styles may gain short-term trust/adoption advantages, creating incentives to optimize explanations for persuasion rather than truthfulness — a moral hazard and potential race-to-the-bottom in explanation fidelity.
- Information asymmetry & signaling:
- Explanations become an endogenous signaling instrument. Without verifiable provenance, consumers cannot easily distinguish truthful from persuasive-but-misleading explanations, exacerbating asymmetric information and potentially increasing demand for third-party verification or reputation mechanisms.
- Platform competition & externalities:
- Platforms that allow (or fail to police) manipulative explanation framings can impose negative externalities across users and markets (fraud, systemic misinformed behavior). Competitors may be forced to match persuasive styles to retain users, amplifying risk.
- Regulation, liability, and insurance:
- Policy responses could include mandatory provenance (signed, auditable explanations), disclosure requirements for explanation sources/uncertainty, and liability rules for demonstrably misleading explanation practices. Insurers and regulators will need new models to price risks linked to cognitive-layer manipulation.
- Design economics and product strategy:
- Firms should internalize costs of user harm from manipulative explanations; invest in verifiable explanation protocols, UI constraints (e.g., standardized uncertainty cues, limited rhetorical authority when confidence is low), and auditing tools measuring trust miscalibration gap as part of deployment risk assessment.
- Research and measurement:
- The trust miscalibration gap can be incorporated into economic models of adoption, platform trustworthiness, and mechanism design (e.g., contracts that reward calibrated transparency). Empirical IO and behavioral-economics work can quantify welfare impacts of persuasive explanations on market outcomes.
- Policy and market interventions:
- Interventions include certification/audits of explanation fidelity, regulation of explanation provenance, platform-level guardrails on persuasive styles, user education, and industry standards for uncertainty signaling.
- Future economic questions:
- How do persuasion-capable AI explanations alter equilibrium adoption and pricing? When do firms have incentives to sell persuasion vs. truthfulness? What liability and monitoring regimes minimize social loss given enforcement costs?
Summary: Treat explanations as an economically meaningful product feature and risk vector. The paper provides both a metric (trust miscalibration gap) and empirical evidence that persuasive LLM explanations can create durable market-relevant harms; policymakers, platform designers, and economists should incorporate cognitive-layer manipulation into models of incentives, regulation, and market design.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. Other | null_result | conceptual definition of a new attack class (AEAs) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Decision Quality | null_result | trust miscalibration gap (difference in human trust between benign and adversarial explanations) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Decision Quality | null_result | experimental manipulation across four explanation-framing dimensions (reasoning mode, evidence type, communication style, presentation format) |
Reading fidelity
high
Study strength
high
|
n=200
|
| Users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. Decision Quality | positive | human-reported trust in AI explanations (trust rating) for benign vs. adversarial explanations |
Reading fidelity
high
Study strength
medium
|
n=200
|
| Persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong (behavioral risk). Decision Quality | positive | preservation of user trust when AI predictions are incorrect |
Reading fidelity
high
Study strength
medium
|
n=200
|
| The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Decision Quality | positive | vulnerability to AEAs measured as preserved trust under expert-like explanation framing |
Reading fidelity
high
Study strength
medium
|
n=200
|
| Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI. Decision Quality | positive | differential vulnerability (preserved trust) across task difficulty, domain type, and participant demographics/attitudes |
Reading fidelity
high
Study strength
medium
|
n=200
|