The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-assisted pulmonary-embolism detection raised radiologist throughput and agreement over time without slowing diagnoses or increasing mortality: agreement with AI-positive flags rose from 70% to 88% over two years while scan volume grew 16% and per-radiologist monthly caseload nearly doubled; clinicians show wide heterogeneity in acceptance rates.

Human-AI Collaboration in Radiology: The Case of Pulmonary Embolism
Paul Goldsmith-Pinkham, Chenhao Tan, Alexander K. Zentefis · January 19, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Paul Goldsmith-Pinkham unresolved corpus identity
  2. Chenhao Tan unresolved corpus identity
  3. Alexander K. Zentefis unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Paul Goldsmith-Pinkham provider ID
  2. Chenhao Tan provider ID
  3. Alexander K. Zentefis provider ID
In a large real-world rollout, radiologists increasingly accept AI flags for pulmonary embolism over two years while maintaining diagnostic speed and patient mortality, with substantial heterogeneity across clinicians and engagement levels.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We study how radiologists use AI to diagnose pulmonary embolism (PE), tracking over 100,000 scans interpreted by nearly 400 radiologists during the staggered rollout of a real-world FDA-approved diagnostic platform in a hospital system. When AI flags PE, radiologists agree 84% of the time; when AI predicts no PE, they agree 97%. Disagreement evolves substantially: radiologists initially reject AI-positive PEs in 30% of cases, dropping to 12% by year two. Despite a 16% increase in scan volume, diagnostic speed remains stable while per-radiologist monthly volumes nearly double, with no change in patient mortality -- suggesting AI improves workflow without compromising outcomes. We document significant heterogeneity in AI collaboration: some radiologists reject AI-flagged PEs half the time while others accept nearly always; female radiologists are 6 percentage points less likely to override AI than male radiologists. Moderate AI engagement is associated with the highest agreement, whereas both low and high engagement show more disagreement. Follow-up imaging reveals that when radiologists override AI to diagnose PE, 54% of subsequent scans show both agreeing on no PE within 30 days.

Summary

Main Finding

In a large real-world deployment of an FDA-cleared AI diagnostic tool for suspected pulmonary embolism (PE), radiologists and AI exhibit systematic, asymmetric, and persistent patterns of agreement and disagreement. AI deployment coincided with higher per-radiologist throughput and stable per-scan diagnostic times and patient mortality, suggesting workflow/productivity gains rather than degraded outcomes. However, substantial heterogeneity across radiologists and evolving disagreement over time imply active human judgment (not uniform automation) governs AI uptake.

Key Points

  • Sample and setting
    • 117,063 CTPA scans interpreted by 389 signing radiologists across a large academic health system; AI rolled out staggeredly across eight sites between Aug 2019 and Jul 2022.
  • Agreement rates (overall, descriptive)
    • When AI predicts no PE: radiologists agree 97% of the time.
    • When AI flags PE: radiologists agree 84% of the time (13 percentage point asymmetry).
  • Time evolution
    • Early post-deployment: radiologists reject AI-positive flags in ~30% of cases.
    • By year two: rejection rate for AI-positive flags falls to ~12% and then stabilizes.
    • Disagreement on AI-negative predictions remains low (2–3%) throughout.
  • Productivity and outcomes
    • Total scan volume up 16% over the period.
    • Average order-to-diagnosis time per scan stable at ~3.6 hours.
    • Per-radiologist monthly volumes increased from 5.3 to 10.4 scans.
    • Patient mortality (30-, 90-day, 1-year) roughly unchanged after rollout.
    • Reading time for AI-positive (PE) cases increased slightly (3.2 → 3.6 hours), consistent with more careful review of flagged cases.
  • Heterogeneity and engagement
    • Wide cross-radiologist variation: some override AI-positive flags ~50% of the time, others rarely do.
    • Female radiologists are ~6 percentage points less likely to override AI-detected PEs than male radiologists.
    • Non-monotonic relationship between AI engagement and agreement when AI flags PE:
      • Moderate engagement (hover ~26% of alerts): highest agreement (91%).
      • Low engagement (hover ~8%) and high engagement (hover ~45%): lower agreement (~81%).
    • Interpretation: low engagement may reflect passive bypass; high engagement may reflect active vetting.
  • Follow-up imaging
    • Of 35,896 patients with initial CTPA and AI, 1,375 (3.8%) had follow-up imaging within 30 days.
    • When radiologists override AI to diagnose PE (i.e., radiologist positive, AI negative), 54% of follow-ups show both agreeing on no PE — implying many initial overrides do not persist on subsequent imaging.
    • When radiologists reject AI-positive flags, follow-up scans do not preserve that original disagreement pattern.

Data & Methods

  • Data sources
    • Hospital OMOP electronic health record data: CTPA orders, radiology reports, patient outcomes, and extensive clinical covariates (demographics, vitals, labs, comorbidities, utilization).
    • AI vendor logs: binary AI predictions (PE detected / no PE), timestamps, prioritization behavior, UI indicators (alerts and heatmaps), and engagement metrics (e.g., “hover” over AI alert).
    • National provider registries: radiologist demographics and characteristics.
  • Outcomes and linkages
    • Radiologist diagnosis extracted from clinical reports via natural language processing; linked to AI prediction and timestamps.
    • Follow-up imaging and outcome windows (30, 90 days, 1 year), mortality and utilization indicators extracted from OMOP.
  • Empirical approach
    • Descriptive longitudinal analysis documenting agreement rates, time trends, and aggregate workflow metrics across the staggered rollout.
    • Heterogeneity analysis across radiologist observables (gender, experience) and behavioral engagement patterns.
    • Follow-up scan analysis to examine persistence/resolution of disagreements.
    • Authors note plans for causal estimation exploiting staggered rollout and quasi-random assignment of patients to radiologists in future work; current results are descriptive.
  • Key limitations (as described / implied)
    • Observational/descriptive: causal effects of AI on diagnostic quality and outcomes not definitively estimated here.
    • Ground-truth ambiguity: primary “label” for evaluation is radiologist report (human judgment), not an independent adjudicated gold standard for every case.
    • Single health system and focus on suspected PE (not incidental PE) may limit generalizability.
    • Follow-up imaging subsample is small (3.8% in 30 days), so persistence conclusions are suggestive.
    • Interface and integration evolved over time; some behavior changes could reflect UI improvements rather than radiologist learning about AI accuracy per se.

Implications for AI Economics

  • Evidence of complementarity, not pure substitution
    • Increased per-radiologist throughput with stable per-scan times and unchanged mortality suggests AI acted largely as workflow/triage complement, enabling higher productivity without clear harm.
    • The increase in reading time for AI-flagged positives indicates human judgment remains crucial and is reallocated toward higher-value review tasks.
  • Value of workflow and UI design
    • AI’s prioritization (moving flagged scans to front of queue) and UI (heatmap preview) likely drive productivity gains; small design choices can materially affect how labor and AI are combined.
  • Trust, calibration, and asymmetric uptake
    • Persistent asymmetry (higher acceptance of AI-negatives than AI-positives) and evolution over time illustrate that perceived cost asymmetries (false negatives vs false positives), trust, and risk aversion shape human-AI integration.
    • Non-monotonic engagement suggests heterogeneous behavioral modes (passive vs active), implying one-size-fits-all deployment or training will not achieve uniform outcomes.
  • Distributional and organizational consequences
    • Large heterogeneity across radiologists implies differential productivity and decision patterns; this could affect performance evaluations, compensation, and staffing decisions.
    • Gender and engagement differences point to behavioral heterogeneity that managers and policymakers may want to monitor and address via training or incentive changes.
  • Welfare and policy considerations
    • The observed pattern—improved throughput without worsening mortality—is promising for welfare gains from medical AI, but the potential for automation bias vs learning remains unresolved; monitoring and safeguards are warranted.
    • Regulators and hospitals should track not only algorithmic accuracy but also human-AI interaction metrics (disagreement rates, follow-up persistence, engagement behaviors).
    • Deployment strategies (staggered rollouts, training modules, UI defaults) matter for long-run adoption and for whether AI creates benign productivity gains or induces harmful deference.
  • Research agenda for AI economics
    • Causal identification: leverage staggered rollout and quasi-random patient assignment to estimate causal effects on diagnostic accuracy, false-positive/negative tradeoffs, treatment rates, and patient outcomes.
    • Welfare and labor-market effects: quantify how AI shifts task composition, returns to radiologist skills, and demand for labor.
    • Mechanisms: disentangle learning from automation bias using richer longitudinal and experimental variation in interface cues, feedback provision, and incentives.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Large, high-frequency observational dataset (>100k scans, ~400 radiologists) and staggered rollout provide credible quasi-experimental leverage and allow tracking of clinical outcomes (mortality) and workflow metrics; however, rollout was not randomized and may correlate with unobserved local changes (training, staffing, case-mix, or simultaneous process changes), so causal claims remain subject to selection and confounding concerns. Methods Rigormedium — Study uses rich longitudinal data, examines agreement dynamics, heterogeneity, and downstream imaging, and leverages rollout timing; but the summary lacks explicit controls, robustness checks, balance tests, or formal event-study estimates to rule out pre-trends, and potential measurement/selection issues (e.g., how AI engagement is defined, whether sicker patients concentrate post-rollout) are not addressed in the description. SampleElectronic records of over 100,000 chest CT scans for pulmonary embolism interpretations performed by nearly 400 radiologists within a single hospital system during the staggered deployment of a commercially cleared (FDA-approved) AI diagnostic platform; includes timestamps, radiologist identifiers, AI flags (PE/no-PE), follow-up imaging within 30 days, and patient mortality outcomes. Themeshuman_ai_collab productivity IdentificationExploits a staggered, real-world rollout of an FDA-approved PE diagnostic AI across a hospital system; identification comes from within-radiologist before/after comparisons and time-series variation (event-study / difference-in-differences style) that compare interpretations and outcomes as clinicians adopt the tool and over time, relying on the timing of deployment to isolate AI effects from other contemporaneous changes. GeneralizabilitySingle hospital system — may not generalize to other hospitals, regions, or health systems, Single clinical task (pulmonary embolism detection on CT) — findings may not apply to other imaging tasks or specialties, Results reflect a specific, FDA-approved AI product and its integration into local workflows; other models or integration approaches may perform differently, Workforce composition, training, and local protocols (radiologist experience, gender mix, staffing changes) may limit transferability, Staggered rollout context (implementation support, incentives) may differ from voluntary or market-driven adoption settings

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We tracked over 100,000 scans interpreted by nearly 400 radiologists during the staggered rollout of a real-world FDA-approved diagnostic platform in a hospital system. Other null_result data_collection_scope (number of scans and radiologists tracked)
Reading fidelity high
Study strength high
n=100000
over 100,000 scans; nearly 400 radiologists
0.8
When AI flags PE, radiologists agree 84% of the time. Decision Quality positive agreement between radiologist and AI when AI flags PE
Reading fidelity high
Study strength high
n=100000
84%
0.8
When AI predicts no PE, radiologists agree 97% of the time. Decision Quality positive agreement between radiologist and AI when AI predicts no PE
Reading fidelity high
Study strength high
n=100000
97%
0.8
Radiologists initially reject AI-positive PEs in 30% of cases, dropping to 12% by year two. Decision Quality negative rate at which radiologists reject AI-positive PE findings over time
Reading fidelity high
Study strength high
n=100000
30% initially; 12% by year two
0.8
Scan volume increased 16% during the study period. Organizational Efficiency positive total scan volume
Reading fidelity high
Study strength high
n=100000
16% increase
0.8
Diagnostic speed remained stable despite the rollout. Task Completion Time null_result diagnostic/reporting speed (time to interpretation)
Reading fidelity high
Study strength medium
n=100000
stable (no change reported)
0.48
Per-radiologist monthly volumes nearly doubled. Organizational Efficiency positive per-radiologist monthly scan volume
Reading fidelity high
Study strength medium
n=400
nearly double
0.48
There was no change in patient mortality during the rollout. Consumer Welfare null_result patient mortality
Reading fidelity high
Study strength medium
n=100000
no change reported
0.48
There is significant heterogeneity in AI collaboration: some radiologists reject AI-flagged PEs half the time while others accept nearly always. Decision Quality mixed variation in individual radiologist agreement/rejection rates with AI
Reading fidelity high
Study strength high
n=400
range from ~50% rejection for some radiologists to near-0% rejection for others
0.8
Female radiologists are 6 percentage points less likely to override AI than male radiologists. Decision Quality negative likelihood of overriding AI recommendation by radiologist gender
Reading fidelity high
Study strength medium
n=400
6 percentage points
0.48
Moderate AI engagement is associated with the highest agreement, whereas both low and high engagement show more disagreement. Decision Quality mixed agreement rate by level of AI engagement
Reading fidelity high
Study strength medium
n=100000
inverted-U relationship (moderate engagement -> highest agreement; low/high -> more disagreement)
0.48
When radiologists override AI to diagnose PE, 54% of subsequent scans show both agreeing on no PE within 30 days. Decision Quality negative follow-up imaging agreement (both no PE) within 30 days after an override-to-diagnose decision
Reading fidelity high
Study strength medium
n=100000
54%
0.48
These patterns (stable speed, higher per-radiologist throughput, no mortality change) suggest AI improves workflow without compromising outcomes. Organizational Efficiency positive inferred effect of AI on workflow efficiency and patient outcomes
Reading fidelity high
Study strength speculative
n=100000
interpretive claim (no single numeric effect size reported)
0.08

Notes