The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A machine-learning alert cut repeat complete blood counts by 15% in a randomized pilot across two hospitals, lowering unnecessary inpatient testing without measurable harms; SmartAlert users averaged 1.54 vs 1.82 CBCs within 52 hours (p < 0.01).

SmartAlert: Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Lab Utilization Reduction
April S. Liang, Fatemeh Amrollahi, Yixing Jiang, Conor K. Corbin, Grace Y. E. Kim, David Mui, Trevor Crowell, Aakash Acharya, Sreedevi Mony, Soumya Punnathanam, Jack McKeown, Margaret Smith, Steven Lin, Arnold Milstein, Kevin Schulman, Jason Hom, Michael A. Pfeffer, Tho D. Pham, David Svec, Weihan Chu, Lisa Shieh, Christopher Sharp, Stephen P. Ma, Jonathan H. Chen · December 04, 2025
arxiv rct high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. April S. Liang unresolved corpus identity
  2. Fatemeh Amrollahi unresolved corpus identity
  3. Yixing Jiang unresolved corpus identity
  4. Conor K. Corbin unresolved corpus identity
  5. Grace Y. E. Kim unresolved corpus identity
  6. David Mui unresolved corpus identity
  7. Trevor Crowell unresolved corpus identity
  8. Aakash Acharya unresolved corpus identity
  9. Sreedevi Mony unresolved corpus identity
  10. Soumya Punnathanam unresolved corpus identity
  11. Jack McKeown unresolved corpus identity
  12. Margaret Smith unresolved corpus identity
  13. Steven Lin unresolved corpus identity
  14. Arnold Milstein unresolved corpus identity
  15. Kevin Schulman unresolved corpus identity
  16. Jason Hom unresolved corpus identity
  17. Michael A. Pfeffer unresolved corpus identity
  18. Tho D. Pham unresolved corpus identity
  19. David Svec unresolved corpus identity
  20. Weihan Chu unresolved corpus identity
  21. Lisa Shieh unresolved corpus identity
  22. Christopher Sharp unresolved corpus identity
  23. Stephen P. Ma unresolved corpus identity
  24. Jonathan H. Chen unresolved corpus identity

Semantic Scholar

Latest observation:

  1. April S. Liang provider ID
  2. Fatemeh Amrollahi provider ID
  3. Yixing Jiang provider ID
  4. Conor K. Corbin provider ID
  5. G. Y. Kim provider ID
  6. David Mui provider ID
  7. Trevor Crowell provider ID
  8. Aakash Acharya provider ID
  9. Sreedevi Mony provider ID
  10. S. Punnathanam provider ID
  11. Jack McKeown provider ID
  12. Margaret Smith provider ID
  13. Steven Lin provider ID
  14. Arnold Milstein provider ID
  15. Kevin Schulman provider ID
  16. Jason Hom provider ID
  17. Michael A. Pfeffer provider ID
  18. Tho D. Pham provider ID
  19. David Svec provider ID
  20. Weihan Chu provider ID
  21. L. Shieh provider ID
  22. Christopher Sharp provider ID
  23. Stephen P. Ma provider ID
  24. Jonathan H. Chen provider ID
A randomized pilot across eight acute-care units found that a machine-learning driven CDS (SmartAlert) reduced repetitive CBC testing by 15% (1.54 vs 1.82 results within 52 hours, p < 0.01) without detectable adverse safety effects.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Repetitive laboratory testing unlikely to yield clinically useful information is a common practice that burdens patients and increases healthcare costs. Education and feedback interventions have limited success, while general test ordering restrictions and electronic alerts impede appropriate clinical care. We introduce and evaluate SmartAlert, a machine learning (ML)-driven clinical decision support (CDS) system integrated into the electronic health record that predicts stable laboratory results to reduce unnecessary repeat testing. This case study describes the implementation process, challenges, and lessons learned from deploying SmartAlert targeting complete blood count (CBC) utilization in a randomized controlled pilot across 9270 admissions in eight acute care units across two hospitals between August 15, 2024, and March 15, 2025. Results show significant decrease in number of CBC results within 52 hours of SmartAlert display (1.54 vs 1.82, p <0.01) without adverse effect on secondary safety outcomes, representing a 15% relative reduction in repetitive testing. Implementation lessons learned include interpretation of probabilistic model predictions in clinical contexts, stakeholder engagement to define acceptable model behavior, governance processes for deploying a complex model in a clinical environment, user interface design considerations, alignment with clinical operational priorities, and the value of qualitative feedback from end users. In conclusion, a machine learning-driven CDS system backed by a deliberate implementation and governance process can provide precision guidance on inpatient laboratory testing to safely reduce unnecessary repetitive testing.

Summary

Main Finding

A machine learning–driven EHR-integrated clinical decision support tool (SmartAlert) reduced repetitive inpatient complete blood count (CBC) testing by 15% (mean CBCs within 52 hours: 1.54 vs 1.82, p < 0.01) in a randomized controlled 7‑month pilot across 9,270 admissions, with no detectable adverse effects on near-term safety outcomes.

Key Points

  • Intervention: SmartAlert uses a probabilistic ML model to predict “clinical stability” of CBC components and presents a targeted interruptive alert at order time offering cancellation or reduced frequency (e.g., alternate-day) rather than a blanket restriction.
  • Effect size: 15.4% relative reduction in repetitive CBCs within 52 hours after alert display (1.54 vs 1.82 results); also significant reduction at 28 hours.
  • Safety: No significant differences in ICU transfers, length of stay, 30-day readmission, or mortality in the pilot.
  • Model performance & clinician trust: Retrospective and silent prospective validation tuned to ~90% PPV (clinically prioritized), with clinician co-design and qualitative interviews informing thresholds and UI.
  • Implementation features: Built on a FHIR-enabled DEPLOYR architecture, Azure serverless functions, Epic OPA alerts, 6‑hour batched predictions to mitigate latency, and governance via CDS Committee + FURM (Fair, Useful, Reliable Models) review.
  • Iterative adjustments: Post-deployment clinician feedback led to excluding patients with recent procedures/transfusions or therapeutic IV heparin from triggering alerts.
  • Scalability claim: Framework is extensible to other lab panels and clinical decisions, but evaluation was limited to general med/surg wards at a single health system.

Data & Methods

  • Setting: Two hospitals (academic tertiary and affiliated community) sharing a single Epic EHR instance; standing lab renewals already required every 72 hours.
  • Study design: Randomized controlled pilot (Aug 15, 2024–Mar 15, 2025) across 8 acute-care units. Admissions randomized at encounter opening via an Epic SmartDataElement: treatment (alert visible) vs control (alert hidden).
  • Sample: 9,270 encounters (4,590 treatment; 4,678 control); 486 alerts displayed in treatment, 460 silently triggered in control.
  • Primary outcome: Mean number of CBC results within 52 hours of alert trigger (chosen to allow alternate-day scheduling).
  • Secondary outcomes: Mean CBCs within 28 hours, ICU transfers within 52 hours, encounter-level ICU transfer rate, length of stay, 30-day readmission, encounter mortality.
  • Statistical analysis: Poisson regression for count outcomes (CBCs), Mann–Whitney U for LOS, Fisher’s exact for binary safety outcomes.
  • Model development/validation: Retrospective tuning on 7 years of admissions to target 90% PPV; silent prospective run prior to deployment confirmed PPV ~88–95% across CBC components.
  • Governance & co-design: Multidisciplinary stakeholder engagement, semi-structured clinician interviews (n=18 pre-deployment; n=9 post-deployment), CDS Committee approval with 90‑day review, and institutional FURM assessment.

Implications for AI Economics

  • Direct cost savings (hospital perspective): Extrapolated 15% reduction in repetitive CBCs translates to ~31,500 fewer tests annually in this system; authors estimate ~$13.3M in institution-specific charges avoided (or ~$4.1M using California median hospital charge). These figures reflect billed charges, not net margins—actual financial benefit depends on payer mix, reimbursement rates, and cost structure.
  • Return on investment (ROI) considerations:
    • Potentially rapid operational ROI from reduced reagent use, phlebotomy labor, processing throughput, and reduced downstream costs from iatrogenic anemia.
    • Upfront and recurring costs include model development, integration (FHIR/Azure/EHR connectors), governance (CDS Committee, FURM reviews), human factors design, and ongoing monitoring/maintenance—these can be material, especially in systems without mature data science or CDS infrastructure.
    • The 30-month development/governance timeline in this case suggests multi-quarter to multi-year time horizons before full ROI realization in similar institutions.
  • Value of precision over blanket policies: ML-enabled, patient-specific guidance avoids blocking necessary tests (unlike blunt restrictions), potentially preserving clinician trust and reducing opportunity costs associated with delayed care; higher PPV (trustworthiness) may lower clinician override rates, increasing realized savings per alert.
  • Scalability and marginal cost structure:
    • Once integrated, marginal cost of applying the model to additional units or similar tests is relatively low (software scaling, minor configuration), improving unit economics as scale increases.
    • However, per-site fixed costs (EHR customization, governance, workflow change management) may limit attractiveness for small hospitals unless offered as a managed/shared service.
  • Incentive alignment and stakeholder impacts:
    • Hospitals capture much of the operational savings; payers may realize downstream savings (fewer complications, shorter stays). Misaligned incentives could slow adoption unless savings are contractually shared or regulators/quality programs incentivize de‑implementation of low‑value care.
    • Reductions in billed charges do not automatically translate to higher operating margins—financial models should account for variable vs fixed cost reductions.
  • Risk and compliance economics:
    • Institutional governance (FURM) and safety monitoring add explicit costs but reduce adoption and regulatory risk; robust evaluation (RCT here) increases payer/hospital confidence and may be necessary to scale procurement.
    • Potential liability concerns (missed diagnosis from suppressed tests) require monitoring and possible reserve procedures; the absence of adverse signals in the pilot is supportive but not definitive.
  • Design choices with economic trade-offs:
    • Prioritizing PPV (high precision) reduces false positive alerts and clinician interruptions but lowers recall—this reduces alert volume and thus immediate cost savings per eligible patient; economic optimization requires balancing trust (adoption) vs total tests avoided.
    • Batched precomputation mitigated latency but introduced infrastructure and compute costs; real-time scoring could raise operational costs and complexity.
  • Market and commercialization opportunities:
    • The approach (ML + FHIR + EHR CDS wrapper) is potentially productizable for multi-site deployment, offering recurring licensing or SaaS revenue models; however, competitive differentiation hinges on validated impact, interoperability, and governance support.
  • Equity and systemic effects:
    • Implementation must monitor for differential performance across patient subgroups; model-induced reductions could unintentionally affect care access or outcomes in marginalized populations, with potential economic and reputational costs.
  • Research/evaluation gaps for economic decision-making:
    • Need for full cost-effectiveness analyses incorporating implementation costs, net revenue changes, downstream clinical outcomes (e.g., anemia-related costs), and payer impacts.
    • Multi-center trials and longer-term follow-up would reduce uncertainty and improve investment cases for broader adoption.

Summary statement for decision-makers: SmartAlert demonstrates clinically meaningful reductions in low-value testing with RCT evidence and no short-term safety signals; economically, it offers potentially large operational savings at scale, but realizing net financial benefit requires careful accounting for non-trivial development/integration and governance costs, alignment of incentives across stakeholders, and continued monitoring to ensure safety and equitable impact.

Assessment

Paper Typerct Evidence Strengthhigh — Causal inference is supported by random assignment of admissions in a sizable (n=9,270) multi-unit pilot, with statistically significant reduction in repeat CBCs and safety outcomes assessed; the RCT design minimizes confounding and supports a causal interpretation of the observed utilization change. Methods Rigorhigh — The study uses a randomized design across multiple acute-care units, prespecifies a clear, measurable primary outcome (CBCs within 52 hours), reports significance and effect sizes, and assesses safety outcomes; however, the paper is a pilot in two hospitals so details on randomization clustering, concealment/blinding, model training/validation and contamination controls are limited in the summary. Sample9270 hospital admissions across eight acute care units in two hospitals between August 15, 2024 and March 15, 2025; intervention targeted inpatient complete blood count (CBC) ordering; further inclusion/exclusion criteria and patient demographics not specified in the summary. Themesproductivity human_ai_collab IdentificationRandomized controlled trial: hospital admissions were randomly assigned to receive the SmartAlert CDS display or usual care, and outcomes (number of CBC results within 52 hours, safety measures) were compared between arms to estimate the causal effect of the intervention. GeneralizabilityConducted in two hospitals within a single health system—may not generalize to other systems or countries, Only acute inpatient settings and eight units were studied—results may not hold in outpatient or specialty settings, Intervention targeted CBC ordering specifically; effects on other tests or broader care processes are unknown, Model and CDS integrated with a particular EHR and local workflows—implementation and effects may differ with other EHRs or clinical processes, Pilot timeframe is limited (7 months) so long-term adoption, learning, and sustainability are uncertain, Potential for unblinded clinicians and Hawthorne or contamination effects that could differ elsewhere

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Repetitive laboratory testing unlikely to yield clinically useful information is a common practice that burdens patients and increases healthcare costs. Consumer Welfare negative burden to patients and healthcare costs from repetitive lab testing
Reading fidelity high
Study strength low
not reported
0.3
Education and feedback interventions have limited success, while general test ordering restrictions and electronic alerts impede appropriate clinical care. Decision Quality negative effectiveness of education/feedback and impact of ordering restrictions and alerts on appropriate clinical care
Reading fidelity high
Study strength low
not reported
0.3
We developed SmartAlert, a machine learning-driven clinical decision support system integrated into the electronic health record that predicts stable laboratory results to reduce unnecessary repeat testing. Other positive prediction of stable laboratory results and intended reduction in repeat testing
Reading fidelity high
Study strength speculative
not reported
0.1
The intervention was evaluated in a randomized controlled pilot across 9,270 admissions in eight acute care units across two hospitals between August 15, 2024, and March 15, 2025. Other neutral study sample and trial setting
Reading fidelity high
Study strength high
n=9270
1.0
SmartAlert produced a significant decrease in the number of CBC results within 52 hours of SmartAlert display (1.54 vs 1.82, p <0.01), representing a 15% relative reduction in repetitive testing. Organizational Efficiency positive number of CBC results within 52 hours of SmartAlert display
Reading fidelity high
Study strength high
n=9270
1.54 vs 1.82, p <0.01 (15% relative reduction)
1.0
The reduction in repetitive testing occurred without adverse effect on secondary safety outcomes. Error Rate null_result secondary safety outcomes (unspecified in the summary)
Reading fidelity high
Study strength medium
n=9270
0.6
Implementation lessons learned include interpretation of probabilistic model predictions in clinical contexts, stakeholder engagement to define acceptable model behavior, governance processes for deploying a complex model in a clinical environment, user interface design considerations, alignment with clinical operational priorities, and the value of qualitative feedback from end users. Governance And Regulation positive implementation/process learning outcomes (interpretability, governance, UI, stakeholder engagement, alignment with operations, qualitative feedback)
Reading fidelity high
Study strength medium
not reported
0.6
A machine learning-driven CDS system backed by a deliberate implementation and governance process can provide precision guidance on inpatient laboratory testing to safely reduce unnecessary repetitive testing. Organizational Efficiency positive precision guidance leading to reduced repetitive inpatient laboratory testing
Reading fidelity high
Study strength medium
n=9270
15% relative reduction (reported)
0.6

Notes