The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Polished explanations from LLMs increase confidence but can mislead: in visual reasoning they suppress users' ability to catch model errors, whereas in logical reasoning they improve performance; exposing uncertainty and deferring unclear cases to humans yields better error recovery in many settings.

The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
Ruth Cohen, Lu Feng, Ayala Bloch, Sarit Kraus · January 31, 2026
arxiv rct high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Ruth Cohen unresolved corpus identity
  2. Lu Feng unresolved corpus identity
  3. Ayala Bloch unresolved corpus identity
  4. Sarit Kraus unresolved corpus identity

Semantic Scholar

Latest observation:

  1. R. Cohen provider ID
  2. Lu Feng provider ID
  3. Ayala Bloch provider ID
  4. Sarit Kraus provider ID
Fluent LLM explanations reliably raise user confidence and reliance but do not consistently improve—and can impair—human-AI team accuracy depending on task modality, with probability displays and selective automation improving error recovery in visual tasks while explanations help in logical reasoning tasks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations using a multi-stage reveal design and between-subjects comparisons. In visual reasoning, LLM explanations increase confidence but do not improve accuracy beyond the AI prediction alone, and substantially suppress users' ability to recover from model errors. Interfaces exposing model uncertainty via predicted probabilities, as well as a selective automation policy that defers uncertain cases to humans, achieve significantly higher accuracy and error recovery than explanation-based interfaces. In contrast, for language-based logical reasoning tasks, LLM explanations yield the highest accuracy and recovery rates, outperforming both expert-written explanations and probability-based support. This divergence reveals that the effectiveness of narrative explanations is strongly task-dependent and mediated by cognitive modality. Our findings demonstrate that commonly used subjective metrics such as trust, confidence, and perceived clarity are poor predictors of human-AI team performance. Rather than treating explanations as a universal solution, we argue for a shift toward interaction designs that prioritize calibrated reliance and effective error recovery over persuasive fluency.

Summary

Main Finding

Fluent LLM explanations systematically increase user confidence and agreement with AI predictions but do not reliably improve—and can worsen—objective human–AI team performance in some tasks. The effect is task-dependent: narrative explanations harmed performance on abstract visual reasoning (RAVEN) by masking AI errors, while they improved performance on language-based logical reasoning (LSAT). Simple probability displays and selective automation often produce better calibrated reliance and higher accuracy than explanation-focused interfaces.

Key Points

  • Persuasion Paradox: Narrative fluency raises subjective trust/confidence without guaranteeing (and sometimes reducing) objective accuracy.
  • Task modality matters:
    • Visual/abstract reasoning (RAVEN): LLM explanations increased confidence but did not raise accuracy beyond the AI prediction alone and substantially reduced users’ ability to recover from model errors.
    • Language/logical reasoning (LSAT): LLM explanations produced the highest accuracy and error recovery, outperforming expert-written rationales and probability displays.
  • Alternative supports:
    • Predicted-probability displays yielded higher error-recovery and matched or exceeded model accuracy in RAVEN.
    • A selective automation policy (auto-accept when the model is confident; defer otherwise) achieved the highest accuracy in RAVEN (post hoc).
  • Subjective metrics (trust, perceived clarity) are poor proxies for team performance; interfaces that feel better to users may still worsen outcomes.
  • Early exposure to a given explanation style can prime future reasoning strategies (order effects).

Data & Methods

  • Overall approach: Controlled human-subject experiments comparing prediction-only, prediction+explanation (LLM), prediction+visual explanation (OS heatmaps), prediction+predicted probabilities, and baseline human-only conditions. Measures: objective accuracy, self-reported confidence/trust/clarity, agreement with correct AI predictions, and error-recovery rate when AI is wrong.
  • Study A — RAVEN multi-stage (within-subject):
    • n = 27 participants, up to 16 RAVEN puzzles.
    • Multi-stage reveal: (1) before AI, (2) after AI prediction, (3) after explanation (LLM or OS).
    • Results: baseline accuracy 37.0%; after prediction 49.8%; after explanation 48.8%. Confidence rose only after explanations (mean 3.66 → 3.81).
    • Stats: Friedman tests and Wilcoxon signed-rank post hoc (significant accuracy gain from Stage 1→2; confidence increase Stage 2→3).
  • Study B — RAVEN between-subjects:
    • n = 100 (20 per condition): Human Only; Prediction Only; Prediction + LLM Explanation; Prediction + OS Heatmap; Prediction + Predicted Probabilities. Model accuracy controlled at 60% (6 correct / 4 incorrect per participant).
    • Key results:
      • Objective accuracy: Human Only 24.6%; Prediction Only 53.5%; LLM 57.0%; OS 55.0%; Probabilities 60.5%; Selective automation (post hoc) 69.5%.
      • Error recovery (when AI wrong): LLM 16.2% (lowest); Probability 37.5% (highest among participant conditions); Selective automation 31.2%.
      • Agreement with correct AI: LLM ~84.2%; Probability ~75.8%.
      • Subjective ratings: predicted-probability interface rated highest for clarity/understanding/trust in RAVEN.
      • Stats: Kruskal–Wallis and Dunn post hoc tests (significant differences across conditions).
  • Study C — LSAT between-subjects:
    • n = 80 (20 per condition): Prediction Only; Prediction + LLM Explanation; Prediction + Expert Explanation; Prediction + Predicted Probabilities. Model accuracy again held at 60%.
    • Key results:
      • Objective accuracy: Prediction Only 48.5%; LLM Explanation 72.5% (exceeds AI solo 60%); Expert Explanation 55.5%; Probability 47.0%.
      • Error recovery: LLM 47.5% (highest); Expert 36.2%; Probability 35.0%; Prediction-only 27.5%.
      • Agreement with correct AI: LLM 89.2% (highest).
      • Subjective ratings: LLM explanations rated highest for clarity and understanding in LSAT.
      • Stats: Kruskal–Wallis and Dunn post hoc tests (significant).
  • Models/tools: CNN-based prediction for RAVEN; Claude 3.7 used for LLM-generated rationales; occlusion sensitivity heatmaps for visual explanations.
  • Definitions:
    • Agreement with correct AI = fraction of trials where participant accepted the AI when it was correct.
    • Error recovery = fraction of trials where participant rejected an incorrect AI prediction and selected the correct answer.

Implications for AI Economics

  1. Evaluation and KPIs

    • Don’t equate trust/UX metrics with performance. Procurement, A/B testing, and ROI analyses must prioritize objective task accuracy, error-recovery rates, and decision-quality metrics rather than user satisfaction alone.
    • Contracts and audits should require benchmarked performance across task modalities and include checks for error masking and overreliance.
  2. Product design and deployment strategy

    • One-size-fits-all explanation policies are economically risky. Firms should tailor explanation formats to task modality:
      • For perceptual/visual/abstract tasks, prefer calibrated uncertainty displays (predicted probabilities) and selective automation thresholds to maximize accuracy and reduce costly error propagation.
      • For language and structured-reasoning tasks, fluent LLM rationales can materially improve accuracy and may justify higher investment in narrative explanation pipelines.
    • Implement selective automation (automatically accept high-confidence model outputs; route uncertain cases to humans) as a cost-effective hybrid policy—improves accuracy and reduces human workload relative to explanation-heavy interfaces.
  3. Labor and organizational impacts

    • Training and onboarding should focus on calibrated reliance—teaching users how to interpret probabilistic signals and detect common failure modes—because explanations alone may not teach error-detection skills.
    • Firms estimating labor displacement or augmentation effects must account for task-dependent gains: in some domains (e.g., legal/logical reasoning), LLM explanations can meaningfully augment human output; in others (e.g., visual inference), they may reduce human oversight quality and increase risk.
  4. Risk management and regulation

    • Regulators and safety frameworks should require disclosure of model uncertainty and evidence of human-AI error recovery performance, not just UX satisfaction metrics.
    • For high-stakes domains, mandate domain- and modality-specific validation (including cross-modal tests) and require demonstration that explanations improve objective outcomes or, at minimum, do not materially degrade them.
  5. Economic modeling and ROI

    • Expected value calculations for deploying explanation-capable systems should incorporate:
      • The probability that an explanation will increase correct acceptance versus mask errors (error recovery and agreement rates).
      • Costs of false acceptances (downstream error costs) and benefits of correctly accepted predictions (productivity gains).
      • Task-dependent parameters; sensitivity analyses should be run separately for visual vs language tasks.
    • Investment in LLM explanation infrastructure is not uniformly cost-effective; cost-benefit depends on whether narrative explanations lead to net increases in accuracy or merely increase perceived trust.
  6. Measurement and future research priorities

    • Deployers should instrument systems to continuously measure decision-level outcomes (accuracy, overrides, downstream costs), not only self-reports.
    • Economic research should estimate macro-level impacts of explanation choices on error externalities (e.g., legal, medical, financial harms) and on labor market signaling (skill premiums for humans able to recover from AI errors).

Summary recommendation: Treat explanations as a tool—not a panacea. Prioritize designs and procurement criteria that improve calibrated reliance and error recovery (probabilities, selective automation, task-specific LLM rationales where empirically validated), and align incentive structures and regulations with objective performance, not subjective persuasiveness.

Assessment

Paper Typerct Evidence Strengthhigh — Causal claims are supported by randomized interventions across three controlled experiments and by direct comparisons of interface designs; internal validity is strong. Limitations arise from laboratory task settings, unspecified participant pools, and potential dependence on particular LLMs and explanation formats, which constrain external validity. Methods Rigorhigh — The paper uses careful experimental controls (multi-stage reveal to disentangle prediction vs explanation effects), multiple tasks spanning cognitive modalities, and appropriate between-subject contrasts including alternative support designs (probabilities, selective automation). Potential weaknesses include unspecified sample recruitment details, possible limited power reporting (not in abstract), and reliance on specific explanation implementations. SampleThree controlled human-subject studies in which participants solved abstract visual reasoning problems (RAVEN matrices) and language-based deductive reasoning problems (LSAT-style); participants were randomly assigned to interface conditions (explanations, probability displays, selective automation, etc.). Detailed sample sizes, recruitment platforms, and demographics are not specified in the abstract. Themeshuman_ai_collab productivity IdentificationRandomized controlled human-subject experiments with between-subject assignment and a multi-stage reveal design that isolates the causal effects of AI predictions versus explanations; treatment arms compare explanation-based interfaces, probability/uncertainty displays, and selective automation (defer-to-human) policies. GeneralizabilityLab-based cognitive tasks (RAVEN, LSAT) may not reflect real-world workplace tasks or domain expertise., Participant pool likely consists of online/non-expert subjects rather than professional knowledge workers., Findings may depend on the specific LLM(s), explanation style, and UI implementations tested., Short-term, single-session interactions may not capture learning/adaptation over time., Cultural, language, and demographic heterogeneity of users not addressed, limiting population generalizability.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy (the 'Persuasion Paradox'). Worker Satisfaction mixed user confidence; task accuracy
Reading fidelity high
Study strength medium
not reported
0.6
We conducted three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems) using a multi-stage reveal design and between-subjects comparisons. Other null_result not an outcome — description of experimental method
Reading fidelity high
Study strength high
not reported
1.0
In visual reasoning tasks (RAVEN matrices), LLM explanations increase participants' confidence but do not improve accuracy beyond the AI prediction alone. Decision Quality mixed task accuracy; user confidence
Reading fidelity high
Study strength medium
not reported
0.6
In visual reasoning, LLM explanations substantially suppress users' ability to recover from model errors. Error Rate negative error recovery / ability to correct model errors
Reading fidelity high
Study strength medium
not reported
0.6
Interfaces that expose model uncertainty via predicted probabilities, and a selective automation policy that defers uncertain cases to humans, achieve significantly higher accuracy and error recovery than explanation-based interfaces (in visual reasoning). Decision Quality positive task accuracy; error recovery
Reading fidelity high
Study strength medium
not reported
0.6
For language-based logical reasoning tasks (LSAT problems), LLM explanations yield the highest accuracy and recovery rates, outperforming both expert-written explanations and probability-based support. Decision Quality positive task accuracy; error recovery
Reading fidelity high
Study strength medium
not reported
0.6
The effectiveness of narrative (LLM) explanations is strongly task-dependent and mediated by cognitive modality (visual vs. language-based tasks). Decision Quality mixed explanation effectiveness as measured by accuracy and recovery across task types
Reading fidelity high
Study strength medium
not reported
0.6
Commonly used subjective metrics such as trust, confidence, and perceived clarity are poor predictors of human-AI team performance. Decision Quality negative predictive validity of subjective metrics for team performance (accuracy/error recovery)
Reading fidelity high
Study strength medium
not reported
0.6
Designs should prioritize calibrated reliance and effective error recovery over persuasive fluency; explanations should not be treated as a universal solution. Organizational Efficiency positive design objective (calibrated reliance and error recovery) — normative recommendation
Reading fidelity high
Study strength speculative
not reported
0.1

Notes