The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A citation-grounded LLM raised dermatologists' accuracy by about 12 percentage points, yet when incorrect answers were accompanied by apparently supporting citations clinicians were far more likely to accept them, creating a grounding-dependent safety risk.

Large language models improve physician accuracy but lead to false reliance
Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker · August 01, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Tirtha Chanda unresolved corpus identity
  2. Christoph Wies unresolved corpus identity
  3. Franziska Schramm unresolved corpus identity
  4. Carina Nogueira Garcia unresolved corpus identity
  5. Nicolas B. Merl unresolved corpus identity
  6. Martin J. Hetz unresolved corpus identity
  7. Jochen S. Utikal unresolved corpus identity
  8. Phillip Tschandl unresolved corpus identity
  9. Cristian Navarrete-Dechent unresolved corpus identity
  10. Alexander Thiem unresolved corpus identity
  11. Jakob N. Kather unresolved corpus identity
  12. Consortium unresolved corpus identity
  13. Titus J. Brinker unresolved corpus identity

Semantic Scholar

Latest observation:

  1. T. Chanda provider ID
  2. C. Wies provider ID
  3. Franziska Schramm provider ID
  4. Carina Nogueira Garcia provider ID
  5. N. Merl provider ID
  6. M. J. Hetz provider ID
  7. J. Utikal provider ID
  8. P. Tschandl provider ID
  9. Cristian Navarrete-Dechent provider ID
  10. Alexander Thiem provider ID
  11. J. Kather provider ID
  12. Consortium provider ID
  13. Titus K Brinker provider ID
A retrieval-augmented, citation-displaying LLM (CORA) improved dermatologists' accuracy from 70.8% to 82.6% in a within-subjects reader study, but citation support also substantially increased clinicians' tendency to accept incorrect model advice.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

Summary

Main Finding

An agentic retrieval-augmented LLM (CORA) raised physician accuracy on dermatology questions (unaided 70.8% → assisted 82.6%) and improved base-model performance—particularly on cases published after models’ training cutoffs—but introduced a grounding-dependent safety risk: when physicians perceived the LLM’s citations as supporting its answer, they were much more likely to accept correct advice and much less likely to resist incorrect advice (a “grounding miscalibration” effect).

Key Points

  • Aggregate performance

    • Physician accuracy rose from 70.8% (95% CI 67.0–74.5) unaided to 82.6% (95% CI 80.0–85.2) with CORA (n=46 physicians; 736 decisions).
    • CORA was non-inferior to five backbone models and often improved accuracy, with larger gains for weaker backbones.
  • Model- & dataset-level effects

    • Two dermatology datasets: DermBenchQA (4,855 multiple-choice items) and DermCaseQA (998 open-ended case-report–derived items published after model cutoffs).
    • Retrieval sufficiency was reached for ~45% of items (2,207 matched in DermBenchQA; 437 matched in DermCaseQA).
    • Retrieval gains were much larger on post-cutoff case reports: e.g., Gemma 3 +22.7 pp, Qwen 2.5 +18 pp, GPT-5 +3.4 pp on DermCaseQA.
  • Citation support predicts correctness and drives behavior

    • CORA answers that had at least one physician-rated supportive citation were far more likely to be correct: 92.9% vs 65.8% (OR 6.76).
    • Physician final answers were correct in 87.7% of responses when ≥1 cited source was judged supportive vs 65.5% when none were (OR 3.75).
  • Grounding miscalibration (safety tradeoff)

    • Adoption of correct LLM advice when the physician was initially incorrect (RAIR): overall 64.3%. When a citation was judged supportive RAIR = 76.9% vs 34.0% when unsupported.
    • Resistance to incorrect LLM advice when the physician was initially correct (RSR): overall 64.6%. When a citation was judged supportive RSR = 34.8% vs 92.0% when unsupported.
    • Concretely: perceived citation support increases adoption of correct advice but greatly weakens resistance to incorrect advice—physicians abandoned correct answers in many cases when an incorrect output appeared citation-supported.
  • Reader study design & user behavior

    • Within-subjects: each physician first answered unaided, then saw CORA’s answer with linked citations and rated per-source support, then recorded a final answer.
    • Citation precision (individual sources judged supportive): ~60.6% of cited sources rated supportive; at least one supportive citation on 80.2% of questions (majority vote).

Data & Methods

  • System: CORA — agentic retrieval-augmented generation with an iterative retrieval loop over a vector DB containing guidelines, textbooks, and case reports; generates answers with linked citations.
  • Datasets:
    • DermBenchQA: 4,855 single-best-answer questions compiled from four dermatology QA benchmarks.
    • DermCaseQA: 998 open-ended questions derived from PubMed-indexed dermatology case reports published after model training cutoffs (contamination-resistant).
  • Matching for evaluation: comparisons restricted to items where the retriever reached sufficiency (DermBenchQA: 2,207; DermCaseQA: 437).
  • Backbone models tested: GPT-5, Llama-4, Qwen 2.5, Mistral Large 2, Gemma 3.
  • Reader study:
    • 46 dermatologists from 21 countries (43 completed residency; mix of experience).
    • 320-item subsample of DermBenchQA; each physician answered 16 items; total 736 physician-question decisions.
    • Per-source citation support rated by physicians; question-level support derived by majority vote (2/3).
  • Statistics: Wilcoxon signed-rank for paired physician accuracy; one-sided Wald non-inferiority tests for model comparisons; odds ratios with clustered logistic regression; bootstrap CIs.

Implications for AI Economics

  • Value proposition and pricing

    • Retrieval augmentation meaningfully increases utility, especially for cheaper/weaker backbone models and for up-to-date information. This implies a strong business case for lower-cost LLM + high-quality retrieval stacks (better price-performance than simply using the largest model).
    • “Explainability” via linked citations is monetizable: provenance increases uptake and perceived value, but must be high-quality to avoid harm.
  • Product design and complementary investments

    • Economic returns depend on investment not only in LLMs but in retrievers, indexed high-quality sources, citation-ranking, and UI that communicates citation confidence. Firms that bundle LLMs with robust retrieval will capture more value.
    • Because citation quality drives both beneficial adoption and harmful deference, there is a premium on investments in retrieval precision, provenance verification, and human-in-the-loop workflows. These are ongoing operating costs (curation, indexing, updates).
  • Regulatory, liability, and insurance considerations

    • Grounding miscalibration creates a liability risk: traceable citations can increase clinician deference to incorrect outputs. Insurers, hospitals, and vendors will need contractual and technical mitigations (audit trails, logging, provenance quality thresholds, explicit uncertainty labels).
    • Regulators may require demonstrable provenance accuracy and post-deployment monitoring; compliance increases deployment costs but may become a market entry barrier favoring incumbents.
  • Market structure and competition

    • The finding that retrieval helps smaller/cheaper models supports market niches: specialized RAG services (vertical retrieval + smaller backbone) can compete effectively with general-purpose giant models, lowering barriers for startups.
    • Differentiation may shift from raw model size to dataset curation, retrieval quality, and provenance UI—areas where incumbents with domain content partnerships can exert advantage.
  • Externalities and public-good roles

    • Misleading citations create negative safety externalities; public investment (or standards bodies) in curated, validated medical knowledge bases could reduce industry-wide risk and lower private compliance costs.
    • Public-payments or subsidies for high-quality domain corpora (clinical guidelines, updated case reports) could improve welfare by reducing grounding miscalibration.
  • Metrics, procurement, and reimbursement

    • Procurement and clinical-efficacy evaluation should go beyond aggregate accuracy to metrics that capture appropriate reliance (RAIR/RSR), provenance precision, and harms from grounding miscalibration.
    • Payers and health systems should consider reimbursement models that reward systems demonstrating both accuracy gains and low rates of harmful reversals; this affects ROI calculations.
  • Research and monitoring priorities (economic implications)

    • Ongoing A/B testing, real-world monitoring, and auditing of citation support and clinician behavior are necessary and constitute recurring operating expenditures.
    • There is value in developing certification/audit services and third-party verifiers for retrieval provenance—new markets for assurance services.

Overall: retrieval-augmented LLMs can increase clinical value at lower model cost, but the economic case requires investment in high-quality retrieval, provenance verification, interface design, and governance to manage grounding-miscalibration risks. Purchasers, regulators, and insurers should demand granular reliance-safety metrics and fund curation infrastructure to align incentives.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The study uses real clinicians, a within-subjects paired design, multiple benchmark and held-out case datasets (including post-training case reports), and appropriate statistical tests, giving credible evidence that CORA can change physician decisions; however, lack of randomization, modest physician sample size (n=46), selection of a sufficiency-restricted item subset, and an artificial vignette/readers-study setting limit causal claims about real-world clinical impact and safety. Methods Rigormedium — Analyses are thorough (non-inferiority tests, clustered logistic models, bootstrap CIs, careful stratification by question type and citation support) and the use of contamination-resistant post-cutoff case reports strengthens claims about retrieval benefits. But the design lacks randomization or counterbalancing (always unaided then assisted), uses a subset of items where retrieval succeeded, and is conducted in a simulated vignette environment rather than in clinical workflow, leaving risks of order effects, selection bias, and limited ecological validity. SampleBenchmark datasets: DermBenchQA (4,855 single-best-answer questions pooled from four dermatology QA benchmarks) and DermCaseQA (998 open-ended questions generated from 893 PubMed-indexed dermatology case reports published after model training cutoffs). Reader study: 46 physicians from 21 countries (43 completed residency; experience ranging from <1 year to ≥10 years) who each answered 16 DermBenchQA items (320-item subsample; 736 physician-question decisions total). Model evaluations included five backbone LLMs (GPT-5, Llama 4, Qwen 2.5, Mistral Large 2, Gemma 3) and reader-study answers used CORA with Llama-4 Scout. Themeshuman_ai_collab adoption IdentificationWithin-subjects pre/post design: each physician answered clinical vignettes unaided, then viewed CORA's retrieval-augmented answer and citations and could revise their response; paired comparisons (Wilcoxon signed-rank, clustered logistic regressions, bootstrap CIs) quantify the change in accuracy attributable to seeing CORA. Model-level comparisons use non-inferiority tests on matched items where the retriever reached sufficiency; citation-support labels derive from physician ratings (majority vote or per-physician). No random assignment or concurrent control group was used. GeneralizabilitySpecialty-limited: evaluated only in dermatology; findings may not transfer to other clinical domains, Simulated vignette/reader-study environment, not real clinical workflows or patient outcomes, Small, non-random physician sample (n=46) with potential selection bias and limited representation of practice settings, Analyses often restricted to items where retrieval reached a prespecified sufficiency threshold (subset may not represent all use cases), Reader-study used a particular retrieval/LLM configuration (CORA with Llama-4 Scout); other architectures or UI designs could change effects, Cross-sectional measurement; no data on long-term learning, habituation, or behavior over repeated use

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Viewing CORA's answer and cited sources increased physicians' mean accuracy from 70.8% unaided to 82.6% assisted, an 11.8-percentage-point improvement. Decision Quality positive Physician answer accuracy
Reading fidelity high
Study strength high
n=46
11.8 percentage points
0.8
Citations judged to support CORA's answer were associated with higher physician final-answer accuracy: 87.7% versus 65.5% when no citation was judged supportive. Decision Quality positive Correctness of physicians' final answers
Reading fidelity high
Study strength medium
n=736
OR 3.75
0.48
When physicians were initially incorrect, supportive citations were associated with substantially more correction to the correct final answer: 64.6% versus 23.9% without supporting citations. Decision Quality positive Correction of initially incorrect physician answers
Reading fidelity high
Study strength medium
n=215
OR 5.79
0.48
Perceived citation support increased physicians' adoption of correct CORA advice from 34.0% to 76.9% when physicians had initially answered incorrectly. Decision Quality positive Adoption of correct AI advice
Reading fidelity high
Study strength medium
n=171
42.9 percentage points
0.48
Perceived citation support reduced physicians' resistance to incorrect CORA advice: relative self-reliance fell from 92.0% without supportive citations to 34.8% with supportive citations. Ai Safety And Ethics negative Resistance to incorrect AI advice
Reading fidelity high
Study strength medium
n=48
57.2 percentage-point decrease
0.48
CORA's retrieval-enabled answers were non-inferior to the corresponding non-retrieval base models on the matched DermBenchQA questions across all five evaluated backbone models. Decision Quality null_result LLM answer accuracy
Reading fidelity high
Study strength high
n=2207
non-inferiority margin of 1 percentage point
0.8
Retrieval produced larger accuracy gains on dermatology case reports published after the evaluated models' training cutoffs, including a 22.7-percentage-point gain for Gemma 3. Decision Quality positive LLM answer accuracy on post-training-cutoff cases
Reading fidelity high
Study strength medium
n=437
22.7 percentage points (37.5% to 60.2%)
0.48
CORA's citations were not uniformly supportive: at least one citation was judged supportive for 80.2% of rated questions, while only 60.6% of individual cited sources were rated as supportive. Ai Safety And Ethics mixed Accuracy and supportiveness of retrieved citations
Reading fidelity high
Study strength medium
n=192
80.2% question-level support rate; 60.6% source-level citation precision
0.48

Notes