The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Explanations help when AI is right but amplify errors when it is wrong: explained recommendations increase accuracy by 6.3pp when correct but reduce it by 4.9pp when incorrect, driving over-reliance; selectively providing explanations could raise healthcare value to $2.59bn annually versus $1.82bn under universal transparency.

When Medical AI Explanations Help and When They Harm
Manshu Khanna, Ziyi Wang, Lijia Wei, Lian Xue · December 09, 2025
arxiv rct high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Manshu Khanna unresolved corpus identity
  2. Ziyi Wang unresolved corpus identity
  3. Lijia Wei unresolved corpus identity
  4. Lian Xue unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Manshu Khanna provider ID
  2. Ziyi Wang provider ID
  3. Lijia Wei provider ID
  4. Lian Xue provider ID
In a randomized experiment, explanations improve clinicians' diagnostic accuracy when the AI is correct but substantially worsen accuracy when the AI is wrong because explanations are equally persuasive regardless of recommendation quality, producing an AI transparency paradox with large welfare implications.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We document a fundamental paradox in AI transparency: explanations improve decisions when algorithms are correct but systematically worsen them when algorithms err. In an experiment with 257 medical students making 3,855 diagnostic decisions, we find explanations increase accuracy by 6.3 percentage points when AI is correct (73% of cases) but decrease it by 4.9 points when incorrect (27% of cases). This asymmetry arises because modern AI systems generate equally persuasive explanations regardless of recommendation quality-physicians cannot distinguish helpful from misleading guidance. We show physicians treat explained AI as 15.2 percentage points more accurate than reality, with over-reliance persisting even for erroneous recommendations. Competent physicians with appropriate uncertainty suffer most from the AI transparency paradox (-12.4pp when AI errs), while overconfident novices benefit most (+9.9pp net). Welfare analysis reveals that selective transparency generates \$2.59 billion in annual healthcare value, 43% more than the \$1.82 billion from mandated universal transparency.

Summary

Main Finding

Explanations for AI medical recommendations create an "AI transparency paradox": when AI is correct (73% of cases) explanations improve physician decisions (+6.3 percentage points accuracy), but when AI is incorrect (27% of cases) explanations worsen decisions (−4.9 pp). Explanations therefore amplify AI influence regardless of correctness, producing net gains that are smaller than implied by correct-case effects alone and creating substantial welfare and distributional concerns for regulation and deployment.

Key Points

  • Core paradox

    • AI accuracy in experiment: 73.4%.
    • Explanations increase accuracy by 6.3 pp when AI is correct (p<0.01).
    • Explanations decrease accuracy by 4.9 pp when AI is incorrect (p<0.01).
    • Net effect (given base AI accuracy) is +3.3 pp, only ~52% of the first-best improvement.
    • The swing between explained-correct and explained-incorrect is about 11.2 pp.
  • Mechanism: symmetric persuasiveness

    • Modern LLMs generate fluent explanations conditional on observations, not ground truth, so explanation quality does not differ systematically between correct and incorrect recommendations.
    • Humans treat the fluency/content of explanations as a signal of correctness, applying a persuasiveness weight (formalized as ψ(e)=exp(λ·q(e))) that inflates posteriors.
    • Observed miscalibration: physicians behave as if explained AI has 88.2% accuracy when AI is actually correct (15.2 pp overestimate) and 79.2% when AI is incorrect (6.2 pp overestimate).
  • Cognitive layers of failure

    • Over-reliance (calibration failure): inflated perceived reliability regardless of true accuracy (~83.7% over-reliance when AI correct; 76.6% when incorrect).
    • Discernment failure (signal processing): explanations reduce the differential revision to correct vs incorrect advice.
    • False confidence (metacognitive failure): explanations increase diagnostic certainty even when AI is wrong (+4.6 pp in confidence).
  • Interaction with advice format

    • Counterintuitively, probabilistic presentation (e.g., “70% B”) amplified harms: when AI erred, explanations decreased accuracy by 6.6 pp with probabilistic format vs 3.1 pp with deterministic format (p<0.05). Probabilities created ambiguity that explanations then resolved into misplaced certainty.
  • Heterogeneity and distributional effects

    • Overconfident novices (low competence, high certainty) benefit most from explanations (net +9.9 pp).
    • Competent, appropriately uncertain physicians suffer most when AI errs (−12.4 pp).
    • Benefit-to-harm ratio: ~4.03 for overconfident novices vs ~0.29 for humble experts—suggesting transparency amplifies skill gaps and rewards unjustified confidence.
  • Policy/welfare numbers (healthcare illustration)

    • Counterfactual annual welfare (conservative assumptions: 500M AI-assisted diagnoses; $11,000 per diagnostic error):
      • Universal mandated transparency: $1.82 billion/year (≈52% of first-best).
      • Confidence-based selective transparency (explain only when AI confidence >85th percentile): $2.42 billion/year (≈70%).
      • Competence-adaptive selective transparency (target low-competence users): $2.59 billion/year (≈75%), 43% higher than universal transparency.

Data & Methods

  • Design

    • Lab-in-the-field experiment with 257 medical students (clinical years 4–6) at a major Chinese teaching hospital.
    • Each participant made 15 incentivized diagnostic/prescription decisions → 3,855 total decisions.
    • 2×2 between-subjects factorial: Explanation (with vs without) × Format (Deterministic vs Probabilistic).
    • AI advice generated by GPT-4o with 73.4% accuracy on the chosen scenarios.
  • Procedure (per scenario)

    • Task 1: Prior belief elicitation — allocate 100 percentage points across 5 treatment options (SSQ = sum of squared probabilities used as an “informativeness score”).
    • Task 2: Second-order belief elicitation — predict how informativeness (SSQ) would change after seeing AI advice (measures anticipated learning/ex-ante trust).
    • Task 3: Posterior belief elicitation — update probabilities after receiving AI advice (revealed trust and actual updating).
    • Quadratic scoring rules used to incentivize truthful probability reporting.
  • Treatments

    • Deterministic: “AI recommends medication B.”
    • Probabilistic: “AI recommends 70% B, 30% A” (top-two probabilities).
    • Explained treatments included a medical-style rationale accompanying the recommendation.
  • Implementation details and robustness

    • Stage I: Baseline ability test (10 MCQs), mean accuracy 61.7% (below AI).
    • Random assignment to conditions (N ≈ 64–65 per cell); preregistered (AEA RCT Registry).
    • Extensive post-experiment questionnaires: CRT, demographics, clinical experience, AI attitudes.
    • Identification of mechanisms via comparison of anticipated vs revealed trust, belief revision magnitudes, and confidence measures.

Implications for AI Economics

  • Regulation and information design

    • The paper challenges a core regulatory assumption: more interpretability/transparency uniformly improves outcomes. Mandated universal explanations (e.g., EU AI Act) can create welfare losses relative to selective policies.
    • Conditional transparency (e.g., based on model confidence or user competence) substantially outperforms universal mandates in the experiment and is administratively feasible.
    • Regulators should require empirical impact evaluations (RCTs) for explanation interfaces rather than blanket interoperability/explainability requirements.
  • Market and adoption effects

    • Explanation provision changes demand and adoption patterns: it disproportionately helps and attracts lower-skill or overconfident users while potentially harming skilled users—leading to selection effects and potentially inefficient sorting across tools and buyers.
    • Firms may have incentives to provide persuasive explanations to increase uptake even when those explanations increase erroneous reliance; this creates moral hazard and externalities that regulators and purchasers must address (procurement standards, liability rules, certification).
  • Product design and signaling

    • Probabilistic outputs are not a guaranteed guardrail against over-reliance; how probabilities are combined with fluent explanations matters critically. Design choices (UI, framing, whether to show explanations) should be evaluated jointly.
    • Competence-adaptive interfaces (tailoring explanation display to user expertise) can capture more welfare but require mechanisms to assess user competence and privacy-respecting personalization.
  • Welfare measurement and insurance/liability

    • Standard benefit calculations for AI adoption that consider only average accuracy overstate gains if explanations systematically amplify errors in a nontrivial share of cases.
    • Insurers, hospitals, and purchasers should incorporate explanation-induced miscalibration into risk models and pricing; malpractice/liability regimes may need to consider whether explanations were provided and how.
  • Research and evaluation priorities

    • Need for broader, domain-specific randomized trials of explanation policies (healthcare, lending, sentencing, hiring) before large-scale regulatory mandates.
    • Research agenda: methods to calibrate explanations to ground-truth reliability (e.g., silence when low confidence or when explanation-quality–to–accuracy mappings are weak), metrics for explanation honesty, and mechanisms that make explanations less uniformly persuasive (introduce calibrated hedging cues, highlight uncertainty about causal links).
    • Monitor dynamic effects as LLMs improve: more fluent/accurate explanations can increase both benefit and harm if persuasiveness remains decoupled from correctness.

Overall, the paper highlights that explanation design is an economically consequential product feature: it alters beliefs, behavior, adoption, and welfare in nontrivial ways. AI economists, policymakers, and system designers should treat explanation provision as a policy lever with benefits and costs that depend on model reliability, user competence, and interface choices—not as an unambiguously welfare-improving default.

Assessment

Paper Typerct Evidence Strengthhigh — A randomized experiment with 257 participants and 3,855 decisions provides clean causal estimates of how explanations affect human decision accuracy and over-reliance; effect heterogeneity and welfare calculations are directly estimated. Limitations remain around external validity (lab vignettes, students, one AI/explanation type) but do not undermine internal causal identification. Methods Rigorhigh — Appropriate experimental design (random assignment of explanation treatment and measured AI correctness), sufficiently large number of observations per subject, and explicit heterogeneity and welfare analysis indicate rigorous methods; however, details on pre-registration, randomization checks, and robustness to alternative explanation formats are not provided in the summary and would matter for full appraisal. Sample257 medical students completing 3,855 diagnostic vignette decisions with AI recommendations (AI correct in 73% of cases, incorrect in 27%); experiment varied presence of explanations accompanying AI recommendations and measured decision accuracy, perceived AI accuracy, and welfare implications. Themeshuman_ai_collab governance IdentificationRandomized controlled experiment manipulating whether AI recommendations were accompanied by explanations; subjects (medical students) made repeated diagnostic decisions with randomized AI correctness status, allowing causal attribution of changes in decision accuracy to the presence of explanations and their interaction with AI correctness. GeneralizabilityParticipants are medical students rather than practicing clinicians—behavior may differ with experienced professionals, Vignette/lab decision setting lacks real-world stakes, time pressure, and contextual information present in clinical practice, Findings are based on a specific AI system and explanation method; other models/explanations may be more or less discernible when wrong, Clinical diagnostic tasks may not generalize to other domains (finance, law, manufacturing) where costs of errors and decision processes differ, Short-term experiment — does not capture learning, repeated interaction effects, or long-run adoption dynamics

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Explanations improve decisions when algorithms are correct but systematically worsen them when algorithms err. Decision Quality mixed decision_quality
Reading fidelity high
Study strength medium
n=257
0.6
In an experiment with 257 medical students making 3,855 diagnostic decisions, explanations increase accuracy by 6.3 percentage points when AI is correct (73% of cases). Decision Quality positive decision_quality
Reading fidelity high
Study strength medium
n=257
6.3 percentage points (when AI is correct); AI correct in 73% of cases
0.6
Explanations decrease accuracy by 4.9 percentage points when the AI is incorrect (27% of cases). Decision Quality negative decision_quality
Reading fidelity high
Study strength medium
n=257
4.9 percentage points decrease (when AI is incorrect); AI incorrect in 27% of cases
0.6
The asymmetry (benefit when correct, harm when wrong) arises because modern AI systems generate equally persuasive explanations regardless of recommendation quality — physicians cannot distinguish helpful from misleading guidance. Decision Quality negative decision_quality
Reading fidelity high
Study strength medium
n=257
0.6
Physicians treat explained AI as 15.2 percentage points more accurate than reality, with over-reliance persisting even for erroneous recommendations. Decision Quality positive decision_quality
Reading fidelity high
Study strength medium
n=257
15.2 percentage points (overestimation of AI accuracy when explanations provided)
0.6
Competent physicians with appropriate uncertainty suffer most from the AI transparency paradox (−12.4 percentage points when AI errs). Decision Quality negative decision_quality
Reading fidelity high
Study strength medium
-12.4 percentage points (subgroup: competent physicians when AI errs)
0.6
Overconfident novices benefit most from explained AI (+9.9 percentage points net). Decision Quality positive decision_quality
Reading fidelity high
Study strength medium
+9.9 percentage points (subgroup: overconfident novices, net effect)
0.6
Selective transparency generates $2.59 billion in annual healthcare value, 43% more than the $1.82 billion from mandated universal transparency. Consumer Welfare positive consumer_welfare
Reading fidelity high
Study strength speculative
$2.59 billion (selective transparency); $1.82 billion (universal transparency); 43% more
0.1
In the experiment, the AI was correct in 73% of cases and incorrect in 27% of cases. Decision Quality null_result decision_quality
Reading fidelity high
Study strength medium
n=3855
73% correct; 27% incorrect
0.6

Notes