1 cumulative citations
View corpus contextAI tools routinely underdeliver on promised productivity gains: developers who expected a 24% speedup were slowed 19%, and clinical tools often save far less time than vendor claims, with some delivering no measurable benefit; integration costs, verification, and uneven distribution of gains explain the shortfall.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. We review controlled trials and independent validations across software engineering, clinical documentation, and clinical decision support to quantify this expectation-realisation gap. In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note, and one widely deployed tool showed no statistically significant effect. In clinical decision support, externally validated performance falls substantially below developer-reported metrics. These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in.
Summary
Main Finding
Pre-deployment expectations for agentic AI systems (autonomous multi-step agents) systematically exceed realised benefits across software engineering, clinical documentation, and clinical decision support. Measured outcomes in deployment-grade evaluations often show much smaller gains, null effects, or even net harms once integration, verification, and heterogeneity are accounted for.
Key Points
- Large expectation–realisation gaps documented:
- Software engineering: experienced developers expected a 24% speedup but were 19% slower when using an AI tool (43 percentage-point calibration error; METR RCT). Contrasting lab/task evidence shows large speedups for constrained tasks (e.g., a GitHub Copilot trial reported +56% speed on a standardized task).
- Clinical documentation: vendor claims of multi-minute savings (e.g., “~5 minutes per encounter”) contrast with measured reductions under one minute per note in real-world studies. UCLA RCT: one scribe (Nabla) reduced time-in-note by 9.5% vs control; a second (DAX) showed no significant effect; cohort and pre/post studies report ~46–57 seconds saved per note, with partial adoption common.
- Clinical decision support: developer-reported performance inflated relative to external validation (Epic Sepsis Model AUROC 0.63 externally vs 0.76–0.83 reported internally). High-profile concordance claims (e.g., Watson) did not generalise (strict concordance ~49% in a study of colon cancer).
- Recurrent mechanistic drivers:
- Workflow integration friction and partial adoption dilute intended gains (often low and persistent uptake).
- Verification and review burden: human oversight (editing, debugging, safety checks) frequently offsets gross time or accuracy gains.
- Measurement construct mismatch: vendor/developer metrics and lab-task measurements often do not correspond to deployment-grade outcomes (e.g., “minutes saved per encounter” vs “time-in-note”; AUROC vs operational utility at chosen thresholds).
- Heterogeneous treatment effects are the norm: benefits concentrate among less-experienced users, inefficient documenters, or rare/high-context tasks. Experienced/highly optimised users may see little benefit or net harm.
- Net effects can include shifted effort (e.g., less time-in-note but more after-hours EHR), new risks (security vulnerabilities in generated code), and degraded durable learning/skills.
Data & Methods
- Evidence synthesised from controlled trials, field experiments, cohort studies, and external validation studies across three domains:
- Software engineering: METR randomized controlled trial (16 experienced developers, 246 real tasks; intention-to-treat) and other field RCTs/experiments including a standardized GitHub Copilot task (Upwork recruits) and corporate field experiments (Microsoft, Accenture) measuring pull request throughput and task completion times. Security analyses of Copilot-generated snippets reported ~25–33% flagged vulnerabilities.
- Clinical documentation: large RCT at UCLA (238 physicians, ~24,000 encounters per arm) comparing commercial ambient scribes (DAX, Nabla) vs usual care; peer-matched cohort and pre/post studies (integrated delivery systems, Abridge) measuring EHR time-in-note, after-hours EHR time, and adoption rates.
- Clinical decision support: large-scale external validations (e.g., Epic Sepsis Model on 38,455 hospitalisations), retrospective concordance studies comparing vendor-reported concordance (Watson) against local tumour board decisions, with subgroup analyses by age and local guideline constraints.
- Outcomes and analytic choices emphasised:
- Intention-to-treat (ITT) vs per-use estimates (ITT often smaller due to non-adoption).
- Metrics aligned to operational decisions (time-in-note, throughput, AUROC at operational thresholds, concordance rates) rather than lab proxies.
- Heterogeneity analyses by baseline efficiency, experience, task complexity, and demographic/contextual factors.
- Mechanistic inference based on observed adoption rates, qualitative reports of verification/editing burden, and cross-study comparisons of lab vs field settings.
Implications for AI Economics
- Valuation & ROI:
- Do not rely on vendor- or lab-reported headline metrics for financial projections. Discount internally reported performance and translate lab metrics into deployment-grade estimates that account for adoption rates, verification costs, and heterogeneity.
- Use ITT-based projections for organisation-level ROI; model both per-use gross benefits and expected human oversight/verification costs to obtain net benefits.
- Account for distributional returns: average gains mask that benefits may accrue mainly to lower-skilled or less-efficient workers. Investment decisions should target high-yield subpopulations to improve ROI.
- Procurement & deployment policy:
- Require external/independent validation in representative settings (same data distribution, workflows, user profiles) before large-scale procurement.
- Pilot with randomized evaluations (or well-designed quasi-experiments) to estimate ITT effects, adoption dynamics, and long-run costs (including after-hours shifts and skills erosion).
- Contractual procurement terms should include measurable outcome commitments (clear units, baseline, and measurement method) and clauses that adjust price/payments based on realised deployment-grade metrics.
- Measurement & governance:
- Align metrics used for procurement with the metrics used in evaluation (same units, same granularity). Avoid measuring only narrow proxies that overstate value.
- Require explicit accounting for human oversight time in cost–benefit analyses and governance frameworks (e.g., incorporate verification labor into total cost of ownership).
- Incorporate heterogeneity modelling into business cases: specify which user segments and task types are expected to benefit, with sensitivity analyses for uptake and effect size variation.
- Labor market & organizational effects:
- Expect reallocation of effort rather than pure elimination of work—document time savings may be shifted rather than reduced; oversight tasks and quality assurance may increase.
- Consider investments in complementary training and workflow redesign to capture potential gains, and include these costs when estimating net productivity improvements.
- Research & macroeconomic forecasting:
- Macro forecasts of AI-driven productivity should incorporate realistic adoption curves, verification burden, measurement mismatches, and heterogeneous returns; otherwise they risk systematic upward bias.
- Policy-makers and economists should demand field-validated effect sizes for agentic systems and prefer ITT estimates for aggregate impact modelling.
- Operational recommendation:
- Adopt structured planning tools (e.g., the Agentic Automation Canvas) that require quantification across time, quality, risk, enablement, and cost; version and archive expectations to enable ex-post accountability and learning.
Overall, economic assessments of agentic AI deployments must move from trusting benchmark or vendor claims to disciplined, deployment-grade measurement that internalises integration costs, oversight labor, adoption dynamics, and heterogeneity of returns.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. Developer Productivity | negative | realised productivity gains versus expected productivity gains |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. Developer Productivity | negative | developer task completion speed / productivity |
Reading fidelity
high
Study strength
medium
|
24% expected speedup; 19% slowdown observed; 43 percentage-point calibration error
|
| In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note. Task Completion Time | negative | time saved per clinical note |
Reading fidelity
high
Study strength
medium
|
less than one minute per note (measured) versus vendor claims of multi-minute savings
|
| One widely deployed clinical documentation tool showed no statistically significant effect. Task Completion Time | null_result | time per clinical note (efficiency of documentation tool) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In clinical decision support, externally validated performance falls substantially below developer-reported metrics. Decision Quality | negative | prediction/performance metrics of clinical decision support systems |
Reading fidelity
high
Study strength
medium
|
externally validated performance substantially below developer-reported metrics
|
| These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. Organizational Efficiency | negative | factors reducing realised AI benefits (workflow friction, verification burden, measurement mismatch, heterogeneous benefit distribution) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in. Governance And Regulation | positive | policy recommendation uptake / planning practice (normative claim) |
Reading fidelity
high
Study strength
speculative
|
not reported
|