The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI tools routinely underdeliver on promised productivity gains: developers who expected a 24% speedup were slowed 19%, and clinical tools often save far less time than vendor claims, with some delivering no measurable benefit; integration costs, verification, and uneven distribution of gains explain the shortfall.

Quantifying the Expectation-Realisation Gap for Agentic AI Systems
Sebastian Lobentanzer · February 23, 2026
arxiv review_meta medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Sebastian Lobentanzer unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Sebastian Lobentanzer provider ID
A cross-domain review of controlled trials and independent validations finds that deployed AI tools typically deliver much smaller productivity gains than expected—and in some cases slow users down—largely because of workflow frictions, verification burdens, and unequal beneficiary effects.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. We review controlled trials and independent validations across software engineering, clinical documentation, and clinical decision support to quantify this expectation-realisation gap. In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note, and one widely deployed tool showed no statistically significant effect. In clinical decision support, externally validated performance falls substantially below developer-reported metrics. These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in.

Summary

Main Finding

Pre-deployment expectations for agentic AI systems (autonomous multi-step agents) systematically exceed realised benefits across software engineering, clinical documentation, and clinical decision support. Measured outcomes in deployment-grade evaluations often show much smaller gains, null effects, or even net harms once integration, verification, and heterogeneity are accounted for.

Key Points

  • Large expectation–realisation gaps documented:
    • Software engineering: experienced developers expected a 24% speedup but were 19% slower when using an AI tool (43 percentage-point calibration error; METR RCT). Contrasting lab/task evidence shows large speedups for constrained tasks (e.g., a GitHub Copilot trial reported +56% speed on a standardized task).
    • Clinical documentation: vendor claims of multi-minute savings (e.g., “~5 minutes per encounter”) contrast with measured reductions under one minute per note in real-world studies. UCLA RCT: one scribe (Nabla) reduced time-in-note by 9.5% vs control; a second (DAX) showed no significant effect; cohort and pre/post studies report ~46–57 seconds saved per note, with partial adoption common.
    • Clinical decision support: developer-reported performance inflated relative to external validation (Epic Sepsis Model AUROC 0.63 externally vs 0.76–0.83 reported internally). High-profile concordance claims (e.g., Watson) did not generalise (strict concordance ~49% in a study of colon cancer).
  • Recurrent mechanistic drivers:
    • Workflow integration friction and partial adoption dilute intended gains (often low and persistent uptake).
    • Verification and review burden: human oversight (editing, debugging, safety checks) frequently offsets gross time or accuracy gains.
    • Measurement construct mismatch: vendor/developer metrics and lab-task measurements often do not correspond to deployment-grade outcomes (e.g., “minutes saved per encounter” vs “time-in-note”; AUROC vs operational utility at chosen thresholds).
  • Heterogeneous treatment effects are the norm: benefits concentrate among less-experienced users, inefficient documenters, or rare/high-context tasks. Experienced/highly optimised users may see little benefit or net harm.
  • Net effects can include shifted effort (e.g., less time-in-note but more after-hours EHR), new risks (security vulnerabilities in generated code), and degraded durable learning/skills.

Data & Methods

  • Evidence synthesised from controlled trials, field experiments, cohort studies, and external validation studies across three domains:
    • Software engineering: METR randomized controlled trial (16 experienced developers, 246 real tasks; intention-to-treat) and other field RCTs/experiments including a standardized GitHub Copilot task (Upwork recruits) and corporate field experiments (Microsoft, Accenture) measuring pull request throughput and task completion times. Security analyses of Copilot-generated snippets reported ~25–33% flagged vulnerabilities.
    • Clinical documentation: large RCT at UCLA (238 physicians, ~24,000 encounters per arm) comparing commercial ambient scribes (DAX, Nabla) vs usual care; peer-matched cohort and pre/post studies (integrated delivery systems, Abridge) measuring EHR time-in-note, after-hours EHR time, and adoption rates.
    • Clinical decision support: large-scale external validations (e.g., Epic Sepsis Model on 38,455 hospitalisations), retrospective concordance studies comparing vendor-reported concordance (Watson) against local tumour board decisions, with subgroup analyses by age and local guideline constraints.
  • Outcomes and analytic choices emphasised:
    • Intention-to-treat (ITT) vs per-use estimates (ITT often smaller due to non-adoption).
    • Metrics aligned to operational decisions (time-in-note, throughput, AUROC at operational thresholds, concordance rates) rather than lab proxies.
    • Heterogeneity analyses by baseline efficiency, experience, task complexity, and demographic/contextual factors.
  • Mechanistic inference based on observed adoption rates, qualitative reports of verification/editing burden, and cross-study comparisons of lab vs field settings.

Implications for AI Economics

  • Valuation & ROI:
    • Do not rely on vendor- or lab-reported headline metrics for financial projections. Discount internally reported performance and translate lab metrics into deployment-grade estimates that account for adoption rates, verification costs, and heterogeneity.
    • Use ITT-based projections for organisation-level ROI; model both per-use gross benefits and expected human oversight/verification costs to obtain net benefits.
    • Account for distributional returns: average gains mask that benefits may accrue mainly to lower-skilled or less-efficient workers. Investment decisions should target high-yield subpopulations to improve ROI.
  • Procurement & deployment policy:
    • Require external/independent validation in representative settings (same data distribution, workflows, user profiles) before large-scale procurement.
    • Pilot with randomized evaluations (or well-designed quasi-experiments) to estimate ITT effects, adoption dynamics, and long-run costs (including after-hours shifts and skills erosion).
    • Contractual procurement terms should include measurable outcome commitments (clear units, baseline, and measurement method) and clauses that adjust price/payments based on realised deployment-grade metrics.
  • Measurement & governance:
    • Align metrics used for procurement with the metrics used in evaluation (same units, same granularity). Avoid measuring only narrow proxies that overstate value.
    • Require explicit accounting for human oversight time in cost–benefit analyses and governance frameworks (e.g., incorporate verification labor into total cost of ownership).
    • Incorporate heterogeneity modelling into business cases: specify which user segments and task types are expected to benefit, with sensitivity analyses for uptake and effect size variation.
  • Labor market & organizational effects:
    • Expect reallocation of effort rather than pure elimination of work—document time savings may be shifted rather than reduced; oversight tasks and quality assurance may increase.
    • Consider investments in complementary training and workflow redesign to capture potential gains, and include these costs when estimating net productivity improvements.
  • Research & macroeconomic forecasting:
    • Macro forecasts of AI-driven productivity should incorporate realistic adoption curves, verification burden, measurement mismatches, and heterogeneous returns; otherwise they risk systematic upward bias.
    • Policy-makers and economists should demand field-validated effect sizes for agentic systems and prefer ITT estimates for aggregate impact modelling.
  • Operational recommendation:
    • Adopt structured planning tools (e.g., the Agentic Automation Canvas) that require quantification across time, quality, risk, enablement, and cost; version and archive expectations to enable ex-post accountability and learning.

Overall, economic assessments of agentic AI deployments must move from trusting benchmark or vendor claims to disciplined, deployment-grade measurement that internalises integration costs, oversight labor, adoption dynamics, and heterogeneity of returns.

Assessment

Paper Typereview_meta Evidence Strengthmedium — The paper aggregates controlled trials and independent validations (which are relatively strong sources), and reports concrete, replicated shortfalls across domains; however, included studies are heterogeneous in design, scope and sample size, vendor metrics and independent measurements are often non-comparable, and the review does not appear to present a unified, pre-registered meta-analytic synthesis addressing publication or selection bias. Methods Rigormedium — The authors draw on controlled trials and external validations and quantify expectation-realisation gaps, but the abstract provides no detail on systematic search criteria, inclusion/exclusion rules, quality grading, statistical pooling or heterogeneity analysis—so while the approach is methodologically sound in principle, transparency and meta-analytic rigor appear limited. SampleA cross-domain set of controlled trials and independent validation studies from software engineering (experienced developers), clinical documentation (time-per-note studies comparing vendor claims to measured time savings, including at least one widely deployed tool), and clinical decision support (externally validated performance metrics vs developer-reported metrics); specific sample sizes, locations, and trial protocols are not reported in the abstract. Themesproductivity human_ai_collab adoption org_design skills_training IdentificationSynthesizes evidence from randomized controlled trials and independent validation studies across domains; where RCTs exist causal identification comes from random assignment, otherwise from quasi-experimental or pre/post comparisons and external benchmark validation of developer/vendor performance claims. GeneralizabilityLimited to specific domains (software development and clinical workflows) and task types (coding speed, note-writing, decision support), Tools evaluated vary by vendor, capability, and version — findings may not apply to different or newer AI models, Clinical results may be specific to particular healthcare systems, EHR integrations, or regulatory contexts (likely high-income settings), Heterogeneous study designs and measurement constructs reduce comparability across studies, Time-sensitive: rapid model and UX improvements could change realised impacts after the reviewed studies

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. Developer Productivity negative realised productivity gains versus expected productivity gains
Reading fidelity high
Study strength medium
not reported
0.24
In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. Developer Productivity negative developer task completion speed / productivity
Reading fidelity high
Study strength medium
24% expected speedup; 19% slowdown observed; 43 percentage-point calibration error
0.24
In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note. Task Completion Time negative time saved per clinical note
Reading fidelity high
Study strength medium
less than one minute per note (measured) versus vendor claims of multi-minute savings
0.24
One widely deployed clinical documentation tool showed no statistically significant effect. Task Completion Time null_result time per clinical note (efficiency of documentation tool)
Reading fidelity high
Study strength medium
not reported
0.24
In clinical decision support, externally validated performance falls substantially below developer-reported metrics. Decision Quality negative prediction/performance metrics of clinical decision support systems
Reading fidelity high
Study strength medium
externally validated performance substantially below developer-reported metrics
0.24
These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. Organizational Efficiency negative factors reducing realised AI benefits (workflow friction, verification burden, measurement mismatch, heterogeneous benefit distribution)
Reading fidelity high
Study strength speculative
not reported
0.04
The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in. Governance And Regulation positive policy recommendation uptake / planning practice (normative claim)
Reading fidelity high
Study strength speculative
not reported
0.04

Notes