0 cumulative citations
View corpus contextA citation-grounded LLM raised dermatologists' accuracy by about 12 percentage points, yet when incorrect answers were accompanied by apparently supporting citations clinicians were far more likely to accept them, creating a grounding-dependent safety risk.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Summary
Main Finding
An agentic retrieval-augmented LLM (CORA) raised physician accuracy on dermatology questions (unaided 70.8% → assisted 82.6%) and improved base-model performance—particularly on cases published after models’ training cutoffs—but introduced a grounding-dependent safety risk: when physicians perceived the LLM’s citations as supporting its answer, they were much more likely to accept correct advice and much less likely to resist incorrect advice (a “grounding miscalibration” effect).
Key Points
-
Aggregate performance
- Physician accuracy rose from 70.8% (95% CI 67.0–74.5) unaided to 82.6% (95% CI 80.0–85.2) with CORA (n=46 physicians; 736 decisions).
- CORA was non-inferior to five backbone models and often improved accuracy, with larger gains for weaker backbones.
-
Model- & dataset-level effects
- Two dermatology datasets: DermBenchQA (4,855 multiple-choice items) and DermCaseQA (998 open-ended case-report–derived items published after model cutoffs).
- Retrieval sufficiency was reached for ~45% of items (2,207 matched in DermBenchQA; 437 matched in DermCaseQA).
- Retrieval gains were much larger on post-cutoff case reports: e.g., Gemma 3 +22.7 pp, Qwen 2.5 +18 pp, GPT-5 +3.4 pp on DermCaseQA.
-
Citation support predicts correctness and drives behavior
- CORA answers that had at least one physician-rated supportive citation were far more likely to be correct: 92.9% vs 65.8% (OR 6.76).
- Physician final answers were correct in 87.7% of responses when ≥1 cited source was judged supportive vs 65.5% when none were (OR 3.75).
-
Grounding miscalibration (safety tradeoff)
- Adoption of correct LLM advice when the physician was initially incorrect (RAIR): overall 64.3%. When a citation was judged supportive RAIR = 76.9% vs 34.0% when unsupported.
- Resistance to incorrect LLM advice when the physician was initially correct (RSR): overall 64.6%. When a citation was judged supportive RSR = 34.8% vs 92.0% when unsupported.
- Concretely: perceived citation support increases adoption of correct advice but greatly weakens resistance to incorrect advice—physicians abandoned correct answers in many cases when an incorrect output appeared citation-supported.
-
Reader study design & user behavior
- Within-subjects: each physician first answered unaided, then saw CORA’s answer with linked citations and rated per-source support, then recorded a final answer.
- Citation precision (individual sources judged supportive): ~60.6% of cited sources rated supportive; at least one supportive citation on 80.2% of questions (majority vote).
Data & Methods
- System: CORA — agentic retrieval-augmented generation with an iterative retrieval loop over a vector DB containing guidelines, textbooks, and case reports; generates answers with linked citations.
- Datasets:
- DermBenchQA: 4,855 single-best-answer questions compiled from four dermatology QA benchmarks.
- DermCaseQA: 998 open-ended questions derived from PubMed-indexed dermatology case reports published after model training cutoffs (contamination-resistant).
- Matching for evaluation: comparisons restricted to items where the retriever reached sufficiency (DermBenchQA: 2,207; DermCaseQA: 437).
- Backbone models tested: GPT-5, Llama-4, Qwen 2.5, Mistral Large 2, Gemma 3.
- Reader study:
- 46 dermatologists from 21 countries (43 completed residency; mix of experience).
- 320-item subsample of DermBenchQA; each physician answered 16 items; total 736 physician-question decisions.
- Per-source citation support rated by physicians; question-level support derived by majority vote (2/3).
- Statistics: Wilcoxon signed-rank for paired physician accuracy; one-sided Wald non-inferiority tests for model comparisons; odds ratios with clustered logistic regression; bootstrap CIs.
Implications for AI Economics
-
Value proposition and pricing
- Retrieval augmentation meaningfully increases utility, especially for cheaper/weaker backbone models and for up-to-date information. This implies a strong business case for lower-cost LLM + high-quality retrieval stacks (better price-performance than simply using the largest model).
- “Explainability” via linked citations is monetizable: provenance increases uptake and perceived value, but must be high-quality to avoid harm.
-
Product design and complementary investments
- Economic returns depend on investment not only in LLMs but in retrievers, indexed high-quality sources, citation-ranking, and UI that communicates citation confidence. Firms that bundle LLMs with robust retrieval will capture more value.
- Because citation quality drives both beneficial adoption and harmful deference, there is a premium on investments in retrieval precision, provenance verification, and human-in-the-loop workflows. These are ongoing operating costs (curation, indexing, updates).
-
Regulatory, liability, and insurance considerations
- Grounding miscalibration creates a liability risk: traceable citations can increase clinician deference to incorrect outputs. Insurers, hospitals, and vendors will need contractual and technical mitigations (audit trails, logging, provenance quality thresholds, explicit uncertainty labels).
- Regulators may require demonstrable provenance accuracy and post-deployment monitoring; compliance increases deployment costs but may become a market entry barrier favoring incumbents.
-
Market structure and competition
- The finding that retrieval helps smaller/cheaper models supports market niches: specialized RAG services (vertical retrieval + smaller backbone) can compete effectively with general-purpose giant models, lowering barriers for startups.
- Differentiation may shift from raw model size to dataset curation, retrieval quality, and provenance UI—areas where incumbents with domain content partnerships can exert advantage.
-
Externalities and public-good roles
- Misleading citations create negative safety externalities; public investment (or standards bodies) in curated, validated medical knowledge bases could reduce industry-wide risk and lower private compliance costs.
- Public-payments or subsidies for high-quality domain corpora (clinical guidelines, updated case reports) could improve welfare by reducing grounding miscalibration.
-
Metrics, procurement, and reimbursement
- Procurement and clinical-efficacy evaluation should go beyond aggregate accuracy to metrics that capture appropriate reliance (RAIR/RSR), provenance precision, and harms from grounding miscalibration.
- Payers and health systems should consider reimbursement models that reward systems demonstrating both accuracy gains and low rates of harmful reversals; this affects ROI calculations.
-
Research and monitoring priorities (economic implications)
- Ongoing A/B testing, real-world monitoring, and auditing of citation support and clinician behavior are necessary and constitute recurring operating expenditures.
- There is value in developing certification/audit services and third-party verifiers for retrieval provenance—new markets for assurance services.
Overall: retrieval-augmented LLMs can increase clinical value at lower model cost, but the economic case requires investment in high-quality retrieval, provenance verification, interface design, and governance to manage grounding-miscalibration risks. Purchasers, regulators, and insurers should demand granular reliance-safety metrics and fund curation infrastructure to align incentives.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Viewing CORA's answer and cited sources increased physicians' mean accuracy from 70.8% unaided to 82.6% assisted, an 11.8-percentage-point improvement. Decision Quality | positive | Physician answer accuracy |
Reading fidelity
high
Study strength
high
|
n=46
11.8 percentage points
|
| Citations judged to support CORA's answer were associated with higher physician final-answer accuracy: 87.7% versus 65.5% when no citation was judged supportive. Decision Quality | positive | Correctness of physicians' final answers |
Reading fidelity
high
Study strength
medium
|
n=736
OR 3.75
|
| When physicians were initially incorrect, supportive citations were associated with substantially more correction to the correct final answer: 64.6% versus 23.9% without supporting citations. Decision Quality | positive | Correction of initially incorrect physician answers |
Reading fidelity
high
Study strength
medium
|
n=215
OR 5.79
|
| Perceived citation support increased physicians' adoption of correct CORA advice from 34.0% to 76.9% when physicians had initially answered incorrectly. Decision Quality | positive | Adoption of correct AI advice |
Reading fidelity
high
Study strength
medium
|
n=171
42.9 percentage points
|
| Perceived citation support reduced physicians' resistance to incorrect CORA advice: relative self-reliance fell from 92.0% without supportive citations to 34.8% with supportive citations. Ai Safety And Ethics | negative | Resistance to incorrect AI advice |
Reading fidelity
high
Study strength
medium
|
n=48
57.2 percentage-point decrease
|
| CORA's retrieval-enabled answers were non-inferior to the corresponding non-retrieval base models on the matched DermBenchQA questions across all five evaluated backbone models. Decision Quality | null_result | LLM answer accuracy |
Reading fidelity
high
Study strength
high
|
n=2207
non-inferiority margin of 1 percentage point
|
| Retrieval produced larger accuracy gains on dermatology case reports published after the evaluated models' training cutoffs, including a 22.7-percentage-point gain for Gemma 3. Decision Quality | positive | LLM answer accuracy on post-training-cutoff cases |
Reading fidelity
high
Study strength
medium
|
n=437
22.7 percentage points (37.5% to 60.2%)
|
| CORA's citations were not uniformly supportive: at least one citation was judged supportive for 80.2% of rated questions, while only 60.6% of individual cited sources were rated as supportive. Ai Safety And Ethics | mixed | Accuracy and supportiveness of retrieved citations |
Reading fidelity
high
Study strength
medium
|
n=192
80.2% question-level support rate; 60.6% source-level citation precision
|