0 cumulative citations
View corpus contextAI-assisted pulmonary-embolism detection raised radiologist throughput and agreement over time without slowing diagnoses or increasing mortality: agreement with AI-positive flags rose from 70% to 88% over two years while scan volume grew 16% and per-radiologist monthly caseload nearly doubled; clinicians show wide heterogeneity in acceptance rates.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We study how radiologists use AI to diagnose pulmonary embolism (PE), tracking over 100,000 scans interpreted by nearly 400 radiologists during the staggered rollout of a real-world FDA-approved diagnostic platform in a hospital system. When AI flags PE, radiologists agree 84% of the time; when AI predicts no PE, they agree 97%. Disagreement evolves substantially: radiologists initially reject AI-positive PEs in 30% of cases, dropping to 12% by year two. Despite a 16% increase in scan volume, diagnostic speed remains stable while per-radiologist monthly volumes nearly double, with no change in patient mortality -- suggesting AI improves workflow without compromising outcomes. We document significant heterogeneity in AI collaboration: some radiologists reject AI-flagged PEs half the time while others accept nearly always; female radiologists are 6 percentage points less likely to override AI than male radiologists. Moderate AI engagement is associated with the highest agreement, whereas both low and high engagement show more disagreement. Follow-up imaging reveals that when radiologists override AI to diagnose PE, 54% of subsequent scans show both agreeing on no PE within 30 days.
Summary
Main Finding
In a large real-world deployment of an FDA-cleared AI diagnostic tool for suspected pulmonary embolism (PE), radiologists and AI exhibit systematic, asymmetric, and persistent patterns of agreement and disagreement. AI deployment coincided with higher per-radiologist throughput and stable per-scan diagnostic times and patient mortality, suggesting workflow/productivity gains rather than degraded outcomes. However, substantial heterogeneity across radiologists and evolving disagreement over time imply active human judgment (not uniform automation) governs AI uptake.
Key Points
- Sample and setting
- 117,063 CTPA scans interpreted by 389 signing radiologists across a large academic health system; AI rolled out staggeredly across eight sites between Aug 2019 and Jul 2022.
- Agreement rates (overall, descriptive)
- When AI predicts no PE: radiologists agree 97% of the time.
- When AI flags PE: radiologists agree 84% of the time (13 percentage point asymmetry).
- Time evolution
- Early post-deployment: radiologists reject AI-positive flags in ~30% of cases.
- By year two: rejection rate for AI-positive flags falls to ~12% and then stabilizes.
- Disagreement on AI-negative predictions remains low (2–3%) throughout.
- Productivity and outcomes
- Total scan volume up 16% over the period.
- Average order-to-diagnosis time per scan stable at ~3.6 hours.
- Per-radiologist monthly volumes increased from 5.3 to 10.4 scans.
- Patient mortality (30-, 90-day, 1-year) roughly unchanged after rollout.
- Reading time for AI-positive (PE) cases increased slightly (3.2 → 3.6 hours), consistent with more careful review of flagged cases.
- Heterogeneity and engagement
- Wide cross-radiologist variation: some override AI-positive flags ~50% of the time, others rarely do.
- Female radiologists are ~6 percentage points less likely to override AI-detected PEs than male radiologists.
- Non-monotonic relationship between AI engagement and agreement when AI flags PE:
- Moderate engagement (hover ~26% of alerts): highest agreement (91%).
- Low engagement (hover ~8%) and high engagement (hover ~45%): lower agreement (~81%).
- Interpretation: low engagement may reflect passive bypass; high engagement may reflect active vetting.
- Follow-up imaging
- Of 35,896 patients with initial CTPA and AI, 1,375 (3.8%) had follow-up imaging within 30 days.
- When radiologists override AI to diagnose PE (i.e., radiologist positive, AI negative), 54% of follow-ups show both agreeing on no PE — implying many initial overrides do not persist on subsequent imaging.
- When radiologists reject AI-positive flags, follow-up scans do not preserve that original disagreement pattern.
Data & Methods
- Data sources
- Hospital OMOP electronic health record data: CTPA orders, radiology reports, patient outcomes, and extensive clinical covariates (demographics, vitals, labs, comorbidities, utilization).
- AI vendor logs: binary AI predictions (PE detected / no PE), timestamps, prioritization behavior, UI indicators (alerts and heatmaps), and engagement metrics (e.g., “hover” over AI alert).
- National provider registries: radiologist demographics and characteristics.
- Outcomes and linkages
- Radiologist diagnosis extracted from clinical reports via natural language processing; linked to AI prediction and timestamps.
- Follow-up imaging and outcome windows (30, 90 days, 1 year), mortality and utilization indicators extracted from OMOP.
- Empirical approach
- Descriptive longitudinal analysis documenting agreement rates, time trends, and aggregate workflow metrics across the staggered rollout.
- Heterogeneity analysis across radiologist observables (gender, experience) and behavioral engagement patterns.
- Follow-up scan analysis to examine persistence/resolution of disagreements.
- Authors note plans for causal estimation exploiting staggered rollout and quasi-random assignment of patients to radiologists in future work; current results are descriptive.
- Key limitations (as described / implied)
- Observational/descriptive: causal effects of AI on diagnostic quality and outcomes not definitively estimated here.
- Ground-truth ambiguity: primary “label” for evaluation is radiologist report (human judgment), not an independent adjudicated gold standard for every case.
- Single health system and focus on suspected PE (not incidental PE) may limit generalizability.
- Follow-up imaging subsample is small (3.8% in 30 days), so persistence conclusions are suggestive.
- Interface and integration evolved over time; some behavior changes could reflect UI improvements rather than radiologist learning about AI accuracy per se.
Implications for AI Economics
- Evidence of complementarity, not pure substitution
- Increased per-radiologist throughput with stable per-scan times and unchanged mortality suggests AI acted largely as workflow/triage complement, enabling higher productivity without clear harm.
- The increase in reading time for AI-flagged positives indicates human judgment remains crucial and is reallocated toward higher-value review tasks.
- Value of workflow and UI design
- AI’s prioritization (moving flagged scans to front of queue) and UI (heatmap preview) likely drive productivity gains; small design choices can materially affect how labor and AI are combined.
- Trust, calibration, and asymmetric uptake
- Persistent asymmetry (higher acceptance of AI-negatives than AI-positives) and evolution over time illustrate that perceived cost asymmetries (false negatives vs false positives), trust, and risk aversion shape human-AI integration.
- Non-monotonic engagement suggests heterogeneous behavioral modes (passive vs active), implying one-size-fits-all deployment or training will not achieve uniform outcomes.
- Distributional and organizational consequences
- Large heterogeneity across radiologists implies differential productivity and decision patterns; this could affect performance evaluations, compensation, and staffing decisions.
- Gender and engagement differences point to behavioral heterogeneity that managers and policymakers may want to monitor and address via training or incentive changes.
- Welfare and policy considerations
- The observed pattern—improved throughput without worsening mortality—is promising for welfare gains from medical AI, but the potential for automation bias vs learning remains unresolved; monitoring and safeguards are warranted.
- Regulators and hospitals should track not only algorithmic accuracy but also human-AI interaction metrics (disagreement rates, follow-up persistence, engagement behaviors).
- Deployment strategies (staggered rollouts, training modules, UI defaults) matter for long-run adoption and for whether AI creates benign productivity gains or induces harmful deference.
- Research agenda for AI economics
- Causal identification: leverage staggered rollout and quasi-random patient assignment to estimate causal effects on diagnostic accuracy, false-positive/negative tradeoffs, treatment rates, and patient outcomes.
- Welfare and labor-market effects: quantify how AI shifts task composition, returns to radiologist skills, and demand for labor.
- Mechanisms: disentangle learning from automation bias using richer longitudinal and experimental variation in interface cues, feedback provision, and incentives.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We tracked over 100,000 scans interpreted by nearly 400 radiologists during the staggered rollout of a real-world FDA-approved diagnostic platform in a hospital system. Other | null_result | data_collection_scope (number of scans and radiologists tracked) |
Reading fidelity
high
Study strength
high
|
n=100000
over 100,000 scans; nearly 400 radiologists
|
| When AI flags PE, radiologists agree 84% of the time. Decision Quality | positive | agreement between radiologist and AI when AI flags PE |
Reading fidelity
high
Study strength
high
|
n=100000
84%
|
| When AI predicts no PE, radiologists agree 97% of the time. Decision Quality | positive | agreement between radiologist and AI when AI predicts no PE |
Reading fidelity
high
Study strength
high
|
n=100000
97%
|
| Radiologists initially reject AI-positive PEs in 30% of cases, dropping to 12% by year two. Decision Quality | negative | rate at which radiologists reject AI-positive PE findings over time |
Reading fidelity
high
Study strength
high
|
n=100000
30% initially; 12% by year two
|
| Scan volume increased 16% during the study period. Organizational Efficiency | positive | total scan volume |
Reading fidelity
high
Study strength
high
|
n=100000
16% increase
|
| Diagnostic speed remained stable despite the rollout. Task Completion Time | null_result | diagnostic/reporting speed (time to interpretation) |
Reading fidelity
high
Study strength
medium
|
n=100000
stable (no change reported)
|
| Per-radiologist monthly volumes nearly doubled. Organizational Efficiency | positive | per-radiologist monthly scan volume |
Reading fidelity
high
Study strength
medium
|
n=400
nearly double
|
| There was no change in patient mortality during the rollout. Consumer Welfare | null_result | patient mortality |
Reading fidelity
high
Study strength
medium
|
n=100000
no change reported
|
| There is significant heterogeneity in AI collaboration: some radiologists reject AI-flagged PEs half the time while others accept nearly always. Decision Quality | mixed | variation in individual radiologist agreement/rejection rates with AI |
Reading fidelity
high
Study strength
high
|
n=400
range from ~50% rejection for some radiologists to near-0% rejection for others
|
| Female radiologists are 6 percentage points less likely to override AI than male radiologists. Decision Quality | negative | likelihood of overriding AI recommendation by radiologist gender |
Reading fidelity
high
Study strength
medium
|
n=400
6 percentage points
|
| Moderate AI engagement is associated with the highest agreement, whereas both low and high engagement show more disagreement. Decision Quality | mixed | agreement rate by level of AI engagement |
Reading fidelity
high
Study strength
medium
|
n=100000
inverted-U relationship (moderate engagement -> highest agreement; low/high -> more disagreement)
|
| When radiologists override AI to diagnose PE, 54% of subsequent scans show both agreeing on no PE within 30 days. Decision Quality | negative | follow-up imaging agreement (both no PE) within 30 days after an override-to-diagnose decision |
Reading fidelity
high
Study strength
medium
|
n=100000
54%
|
| These patterns (stable speed, higher per-radiologist throughput, no mortality change) suggest AI improves workflow without compromising outcomes. Organizational Efficiency | positive | inferred effect of AI on workflow efficiency and patient outcomes |
Reading fidelity
high
Study strength
speculative
|
n=100000
interpretive claim (no single numeric effect size reported)
|