0 cumulative citations
View corpus contextPolished explanations from LLMs increase confidence but can mislead: in visual reasoning they suppress users' ability to catch model errors, whereas in logical reasoning they improve performance; exposing uncertainty and deferring unclear cases to humans yields better error recovery in many settings.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
While natural-language explanations from large language models (LLMs) are widely adopted to improve transparency and trust, their impact on objective human-AI team performance remains poorly understood. We identify a Persuasion Paradox: fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy. Across three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems), we disentangle the effects of AI predictions and explanations using a multi-stage reveal design and between-subjects comparisons. In visual reasoning, LLM explanations increase confidence but do not improve accuracy beyond the AI prediction alone, and substantially suppress users' ability to recover from model errors. Interfaces exposing model uncertainty via predicted probabilities, as well as a selective automation policy that defers uncertain cases to humans, achieve significantly higher accuracy and error recovery than explanation-based interfaces. In contrast, for language-based logical reasoning tasks, LLM explanations yield the highest accuracy and recovery rates, outperforming both expert-written explanations and probability-based support. This divergence reveals that the effectiveness of narrative explanations is strongly task-dependent and mediated by cognitive modality. Our findings demonstrate that commonly used subjective metrics such as trust, confidence, and perceived clarity are poor predictors of human-AI team performance. Rather than treating explanations as a universal solution, we argue for a shift toward interaction designs that prioritize calibrated reliance and effective error recovery over persuasive fluency.
Summary
Main Finding
Fluent LLM explanations systematically increase user confidence and agreement with AI predictions but do not reliably improve—and can worsen—objective human–AI team performance in some tasks. The effect is task-dependent: narrative explanations harmed performance on abstract visual reasoning (RAVEN) by masking AI errors, while they improved performance on language-based logical reasoning (LSAT). Simple probability displays and selective automation often produce better calibrated reliance and higher accuracy than explanation-focused interfaces.
Key Points
- Persuasion Paradox: Narrative fluency raises subjective trust/confidence without guaranteeing (and sometimes reducing) objective accuracy.
- Task modality matters:
- Visual/abstract reasoning (RAVEN): LLM explanations increased confidence but did not raise accuracy beyond the AI prediction alone and substantially reduced users’ ability to recover from model errors.
- Language/logical reasoning (LSAT): LLM explanations produced the highest accuracy and error recovery, outperforming expert-written rationales and probability displays.
- Alternative supports:
- Predicted-probability displays yielded higher error-recovery and matched or exceeded model accuracy in RAVEN.
- A selective automation policy (auto-accept when the model is confident; defer otherwise) achieved the highest accuracy in RAVEN (post hoc).
- Subjective metrics (trust, perceived clarity) are poor proxies for team performance; interfaces that feel better to users may still worsen outcomes.
- Early exposure to a given explanation style can prime future reasoning strategies (order effects).
Data & Methods
- Overall approach: Controlled human-subject experiments comparing prediction-only, prediction+explanation (LLM), prediction+visual explanation (OS heatmaps), prediction+predicted probabilities, and baseline human-only conditions. Measures: objective accuracy, self-reported confidence/trust/clarity, agreement with correct AI predictions, and error-recovery rate when AI is wrong.
- Study A — RAVEN multi-stage (within-subject):
- n = 27 participants, up to 16 RAVEN puzzles.
- Multi-stage reveal: (1) before AI, (2) after AI prediction, (3) after explanation (LLM or OS).
- Results: baseline accuracy 37.0%; after prediction 49.8%; after explanation 48.8%. Confidence rose only after explanations (mean 3.66 → 3.81).
- Stats: Friedman tests and Wilcoxon signed-rank post hoc (significant accuracy gain from Stage 1→2; confidence increase Stage 2→3).
- Study B — RAVEN between-subjects:
- n = 100 (20 per condition): Human Only; Prediction Only; Prediction + LLM Explanation; Prediction + OS Heatmap; Prediction + Predicted Probabilities. Model accuracy controlled at 60% (6 correct / 4 incorrect per participant).
- Key results:
- Objective accuracy: Human Only 24.6%; Prediction Only 53.5%; LLM 57.0%; OS 55.0%; Probabilities 60.5%; Selective automation (post hoc) 69.5%.
- Error recovery (when AI wrong): LLM 16.2% (lowest); Probability 37.5% (highest among participant conditions); Selective automation 31.2%.
- Agreement with correct AI: LLM ~84.2%; Probability ~75.8%.
- Subjective ratings: predicted-probability interface rated highest for clarity/understanding/trust in RAVEN.
- Stats: Kruskal–Wallis and Dunn post hoc tests (significant differences across conditions).
- Study C — LSAT between-subjects:
- n = 80 (20 per condition): Prediction Only; Prediction + LLM Explanation; Prediction + Expert Explanation; Prediction + Predicted Probabilities. Model accuracy again held at 60%.
- Key results:
- Objective accuracy: Prediction Only 48.5%; LLM Explanation 72.5% (exceeds AI solo 60%); Expert Explanation 55.5%; Probability 47.0%.
- Error recovery: LLM 47.5% (highest); Expert 36.2%; Probability 35.0%; Prediction-only 27.5%.
- Agreement with correct AI: LLM 89.2% (highest).
- Subjective ratings: LLM explanations rated highest for clarity and understanding in LSAT.
- Stats: Kruskal–Wallis and Dunn post hoc tests (significant).
- Models/tools: CNN-based prediction for RAVEN; Claude 3.7 used for LLM-generated rationales; occlusion sensitivity heatmaps for visual explanations.
- Definitions:
- Agreement with correct AI = fraction of trials where participant accepted the AI when it was correct.
- Error recovery = fraction of trials where participant rejected an incorrect AI prediction and selected the correct answer.
Implications for AI Economics
-
Evaluation and KPIs
- Don’t equate trust/UX metrics with performance. Procurement, A/B testing, and ROI analyses must prioritize objective task accuracy, error-recovery rates, and decision-quality metrics rather than user satisfaction alone.
- Contracts and audits should require benchmarked performance across task modalities and include checks for error masking and overreliance.
-
Product design and deployment strategy
- One-size-fits-all explanation policies are economically risky. Firms should tailor explanation formats to task modality:
- For perceptual/visual/abstract tasks, prefer calibrated uncertainty displays (predicted probabilities) and selective automation thresholds to maximize accuracy and reduce costly error propagation.
- For language and structured-reasoning tasks, fluent LLM rationales can materially improve accuracy and may justify higher investment in narrative explanation pipelines.
- Implement selective automation (automatically accept high-confidence model outputs; route uncertain cases to humans) as a cost-effective hybrid policy—improves accuracy and reduces human workload relative to explanation-heavy interfaces.
- One-size-fits-all explanation policies are economically risky. Firms should tailor explanation formats to task modality:
-
Labor and organizational impacts
- Training and onboarding should focus on calibrated reliance—teaching users how to interpret probabilistic signals and detect common failure modes—because explanations alone may not teach error-detection skills.
- Firms estimating labor displacement or augmentation effects must account for task-dependent gains: in some domains (e.g., legal/logical reasoning), LLM explanations can meaningfully augment human output; in others (e.g., visual inference), they may reduce human oversight quality and increase risk.
-
Risk management and regulation
- Regulators and safety frameworks should require disclosure of model uncertainty and evidence of human-AI error recovery performance, not just UX satisfaction metrics.
- For high-stakes domains, mandate domain- and modality-specific validation (including cross-modal tests) and require demonstration that explanations improve objective outcomes or, at minimum, do not materially degrade them.
-
Economic modeling and ROI
- Expected value calculations for deploying explanation-capable systems should incorporate:
- The probability that an explanation will increase correct acceptance versus mask errors (error recovery and agreement rates).
- Costs of false acceptances (downstream error costs) and benefits of correctly accepted predictions (productivity gains).
- Task-dependent parameters; sensitivity analyses should be run separately for visual vs language tasks.
- Investment in LLM explanation infrastructure is not uniformly cost-effective; cost-benefit depends on whether narrative explanations lead to net increases in accuracy or merely increase perceived trust.
- Expected value calculations for deploying explanation-capable systems should incorporate:
-
Measurement and future research priorities
- Deployers should instrument systems to continuously measure decision-level outcomes (accuracy, overrides, downstream costs), not only self-reports.
- Economic research should estimate macro-level impacts of explanation choices on error externalities (e.g., legal, medical, financial harms) and on labor market signaling (skill premiums for humans able to recover from AI errors).
Summary recommendation: Treat explanations as a tool—not a panacea. Prioritize designs and procurement criteria that improve calibrated reliance and error recovery (probabilities, selective automation, task-specific LLM rationales where empirically validated), and align incentive structures and regulations with objective performance, not subjective persuasiveness.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Fluent explanations systematically increase user confidence and reliance on AI without reliably improving, and in some cases undermining, task accuracy (the 'Persuasion Paradox'). Worker Satisfaction | mixed | user confidence; task accuracy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We conducted three controlled human-subject studies spanning abstract visual reasoning (RAVEN matrices) and deductive logical reasoning (LSAT problems) using a multi-stage reveal design and between-subjects comparisons. Other | null_result | not an outcome — description of experimental method |
Reading fidelity
high
Study strength
high
|
not reported
|
| In visual reasoning tasks (RAVEN matrices), LLM explanations increase participants' confidence but do not improve accuracy beyond the AI prediction alone. Decision Quality | mixed | task accuracy; user confidence |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In visual reasoning, LLM explanations substantially suppress users' ability to recover from model errors. Error Rate | negative | error recovery / ability to correct model errors |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Interfaces that expose model uncertainty via predicted probabilities, and a selective automation policy that defers uncertain cases to humans, achieve significantly higher accuracy and error recovery than explanation-based interfaces (in visual reasoning). Decision Quality | positive | task accuracy; error recovery |
Reading fidelity
high
Study strength
medium
|
not reported
|
| For language-based logical reasoning tasks (LSAT problems), LLM explanations yield the highest accuracy and recovery rates, outperforming both expert-written explanations and probability-based support. Decision Quality | positive | task accuracy; error recovery |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The effectiveness of narrative (LLM) explanations is strongly task-dependent and mediated by cognitive modality (visual vs. language-based tasks). Decision Quality | mixed | explanation effectiveness as measured by accuracy and recovery across task types |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Commonly used subjective metrics such as trust, confidence, and perceived clarity are poor predictors of human-AI team performance. Decision Quality | negative | predictive validity of subjective metrics for team performance (accuracy/error recovery) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Designs should prioritize calibrated reliance and effective error recovery over persuasive fluency; explanations should not be treated as a universal solution. Organizational Efficiency | positive | design objective (calibrated reliance and error recovery) — normative recommendation |
Reading fidelity
high
Study strength
speculative
|
not reported
|