Expectation–Realisation Gap for Agentic AI (2026)¶
Source
Document: "Quantifying the Expectation–Realisation Gap for Agentic AI Systems". Author: Sebastian Lobentanzer (Institute of Computational Biology, Helmholtz Munich). Date: February 2026 (arXiv:2602.20292, v2). Distilled: 2026-07-27.
Summary¶
A narrative review that assembles controlled trials and independent external validations to measure the gap between pre-deployment expectations for agentic AI and realised outcomes. Across software engineering, clinical documentation, and clinical decision support, it finds that realised benefits systematically fall short of what users forecast, what vendors claim, and what developers report from internal validation. The headline case is a randomised trial in which experienced developers expected a 24% speedup from AI tools but were slowed by 19%, a 43 percentage-point calibration error that also reversed direction. The review attributes the gap to four recurrent drivers: workflow integration friction and partial adoption, human verification and review burden, measurement construct mismatch, and treatment-effect heterogeneity. It argues for structured planning that states quantified, net-of-oversight benefit expectations up front and ties post-deployment measurement back to them, and proposes the author's own Agentic Automation Canvas as one instrument for doing so.
Hooks, with citations¶
1. The expectation–realisation gap is the review's organising claim. The document: "realised outcomes frequently fall short of pre-deployment expectations. We term this discrepancy the expectation–realisation gap. It reflects systematic patterns in how agentic systems interact with human workflows, how performance is measured, and how benefits are distributed across user populations" (Introduction). Relevance: names a planning problem that sits upstream of tool choice and evaluation; bears on how agentic deployments are planned and governed. Audiences: practitioners, providers, governance.
2. The sharpest single data point: experts forecast a speedup and were slowed. The document: "participants forecast that AI assistance would reduce their completion time by 24%. The measured outcome was a 19% increase in completion time—a 43 percentage-point calibration error on the time-change scale, and a complete reversal in direction." "AI-assisted developers also estimated 20% reduction in completion time after they had performed the task; the opposite of what had happened." "Economics experts (N=34) and machine learning experts (N=54) overestimated the expected speedup even more dramatically, with 39% and 38%, respectively" (Evidence from controlled trials, software engineering copilots; METR RCT, 16 developers, 246 tasks). Relevance: quantifies miscalibration that persists even after the task is done, evidence that self-forecasts and self-reports are unreliable inputs to planning. Audiences: practitioners, governance.
3. On constrained tasks the same class of tool helps, and the benefit is uneven. The document: treated participants "completed a standardised, self-contained programming task 56% faster (95% CI 21–89%)", their self-estimated gains "averaged approximately 35%, meaning they underestimated the realised speedup", and "the benefit was greater for less experienced developers." "These findings are the first instance of heterogeneous treatment effects of agentic AI" (software engineering copilots; GitHub Copilot RCT, Upwork recruits). Relevance: gains are real in specific, well-defined contexts and vary by task complexity and user experience; supports matching method to task (BP-01). Audiences: practitioners, providers.
4. Throughput gains do not account for the cost of verifying output. The document: developers assigned to Copilot completed "12.9–21.8% more pull requests per week at Microsoft and 7.5–8.7% more at Accenture", but "independent security analyses find that 32.8% of Python and 24.5% of JavaScript snippets generated by Copilot are flagged with security issues", and "Copilot can replicate known-vulnerable code patterns at rates around 33%." "Productivity gains that increase time for human oversight (such as review, remediation, and incident risk) are not net gains" (software engineering copilots). Relevance: gross output metrics omit review and remediation cost; bears directly on human oversight (BP-09). Audiences: practitioners, providers, governance.
5. Vendor time-savings claims contrast with measured minutes, and one tool showed no effect. The document: Microsoft publicised "5 minutes saved per clinician per encounter on average" for DAX Copilot; a UCLA RCT across "238 physicians in 14 specialties" found "Nabla reduced time-in-note by 9.5% relative to control (95% CI −17.2 to −1.8; P=0.02), while DAX showed no statistically significant effect (−1.7%, 95% CI −9.4 to +5.9; P=0.66)"; the tools "were used in only approximately 30–34% of visits, and roughly 15% of treatment-group physicians never used their assigned scribe at all" (clinical documentation agents). Relevance: marketing figures and deployment-grade measurement diverge, and partial adoption is a persistent contextual factor, not a temporary onboarding issue. Audiences: practitioners, providers, governance.
6. Documentation time can shift rather than fall. The document: a DAX cohort found "documentation EHR time fell from 5.3 to 4.5 minutes per patient—a saving of approximately 46 seconds—while after-hours EHR time worsened significantly, suggesting time-shifting rather than uniform savings" (clinical documentation agents). Relevance: a headline time saving can mask a redistribution of effort; net benefit needs the whole workflow, not one slice. Audiences: practitioners, governance.
7. Perceived benefit does not track measured benefit. The document: in a study of 252 physicians, "86.5% perceived that their documentation time had decreased, yet there was no overall association between perceived reductions and objectively measured time changes (OR 0.975, P=0.144)"; the objective effect was modest, "approximately 30 seconds lower documentation time per scheduled hour" per 10-point usage increase (clinical documentation agents). Relevance: user satisfaction is not evidence of realised gain; reinforces measuring outcomes rather than trusting perception (BP-08). Audiences: practitioners, governance.
8. Developer-reported accuracy exceeds externally validated accuracy. The document: the Epic Sepsis Model was "externally validated in a large academic health system (38 455 hospitalisations) with an area under the receiver operating characteristic curve (AUROC) of 0.63 (95% CI 0.62–0.64), while Epic previously reported AUROC values of 0.76–0.83", achieving "only 33% sensitivity (missing two thirds of septic patients)"; IBM "publicised concordance rates as high as 96% for Watson for Oncology", while a Korean retrospective study "found strict concordance of 48.9% for colon cancer" (clinical decision support). Relevance: internally reported metrics set expectations that external, deployment-grade evaluation does not reproduce; the core case for adopter-side evaluation (BP-08). Audiences: practitioners, providers, governance.
9. Three mechanistic drivers explain why expectations overshoot. The document: the gap is driven by "workflow integration friction and partial adoption", "verification and review burden", and "measurement construct mismatch", "consistent with planning fallacies widely reported in the psychology literature." "In the METR software engineering trial, the net slowdown occurred because the time spent reviewing, debugging, and integrating AI-generated code exceeded the time saved in initial generation" (Why expectations overshoot). Relevance: friction and partial adoption bear on interface design (BP-06); verification burden on oversight (BP-09); construct mismatch on evaluation (BP-08). Audiences: providers, practitioners, governance.
10. Short-horizon metrics miss durable costs to skill and understanding. The document: an RCT in higher education found students using ChatGPT as a study aid "scored significantly lower on a surprise retention test 45 days later (57.5% vs 68.5%, Cohen's d = 0.68)", and a parallel software engineering study found "AI use impaired the users' conceptual understanding, code reading, and debugging abilities, without delivering significant efficiency gains on average" (Why expectations overshoot). Relevance: benefit projected over short windows can invert over longer ones through cognitive offloading and skill erosion; a planning consideration for research training. Audiences: practitioners, governance.
11. Heterogeneity is the default; there is no stable average gain to plan against. The document: "there is currently no stable, globally positive treatment effect for agentic AI. Average headline figures ... will systematically misrepresent the benefit realised by any specific user, team, or organisation." A field study of "5,172 agents found an average 15% productivity increase" concentrated "among less experienced and lower-skilled workers, while the most experienced agents saw smaller gains and occasional quality declines"; the METR trial "specifically selected experienced developers ... and this is the population that was slowed" (Heterogeneity as the default). Relevance: planning on an average expected gain over-invests in low-yield deployments and under-invests in targeted high-yield ones; who benefits must be modelled, not assumed. Audiences: practitioners, providers, governance.
12. The review prescribes structured planning with quantified, net-of-oversight expectations. The document: "benefit expectations must be explicit and quantified across all relevant dimensions"; "expectations should capture dual perspectives" (user expectation and developer feasibility); "human oversight costs must be deducted"; "outcome metrics must link back to initial expectations in the same units and at the same level of granularity"; "heterogeneity should be modelled explicitly." These are operationalised by "the Agentic Automation Canvas (AAC)", which "captures user expectations as quantified benefit metrics across five dimensions—time, quality, risk, enablement, and cost—with baseline values, confidence levels ... and explicit accounting for human oversight" (Implications for structured planning). Relevance: motivates a planning practice upstream of the operational practices; the AAC is the author's own instrument (see cautions). Audiences: practitioners, providers, governance.
Mapping to practices¶
| Hook | Supports | In tension with |
|---|---|---|
| 1 | BP-01 | |
| 2 | BP-08 | |
| 3 | BP-01 | |
| 4 | BP-09 | |
| 5 | BP-08 | |
| 6 | BP-09 | |
| 7 | BP-08 | |
| 8 | BP-08 | |
| 9 | BP-06, BP-09, BP-08 | |
| 10 | BP-09 | |
| 11 | BP-01 | |
| 12 | BP-01 |
No direct tension with an existing practice. The review sits upstream of the record: it argues that the expected net benefit of an agentic method should be quantified and tested before it is committed to. That planning step is an extension of choosing the method for the task (BP-01), so it is proposed as a note on that page rather than a new practice.
Proposed changes to practices¶
To follow the anti-bloat rule (merge rather than add), the planning point is woven into BP-01 rather than made a new practice, and this review is cited as downweighted external context throughout, not as a driver of change. The author is a contributor to this record, so the document should not carry more weight than the independent evidence it aggregates.
- [x] Add a "what it looks like in practice" note on BP-01, with no change to the practice statement wording: choosing a method includes deciding whether to adopt at all, so state the net benefit you expect before committing, quantified against a baseline and with human oversight cost deducted, and do not assume an average headline gain applies to your users; re-measure the realised outcome against that expectation. This extends "you cannot pick the right method without first understanding the task" to include understanding the expected net benefit. Applied 2026-07-27 (Reasons and Examples note).
- [x] Add this review as a supporting/context source to BP-01: treatment effects are heterogeneous and task-dependent (constrained tasks sped up, expert high-context work slowed), so method fit is decided per task and per user population, not by an average gain. Atoms bp1-a1, bp1-a3. Applied 2026-07-27.
- [ ] Optional grounding as context, if the editors want it, kept secondary to the existing sources on each page:
- BP-08: developer-reported and vendor-reported metrics systematically exceed externally validated ones (Epic Sepsis AUROC 0.76–0.83 reported vs 0.63 validated; Watson 96% vs 48.9% strict concordance; perceived vs measured documentation time). Atoms bp8-a2, bp8-a3.
- BP-09: verification and review cost must be deducted from projected benefit; "productivity gains that increase time for human oversight ... are not net gains", and verification burden drove the METR net slowdown. Atom bp9-a5.
- BP-06: workflow integration friction and partial adoption are leading drivers of the realised-benefit shortfall. Atom bp6-a1.
- [x] Add provenance edges under the named atoms with
ref: expectation-realisation-gap-2026,stance: supports, and the locators/quotes from the hooks. Cited as external field evidence, consistent with how the other out-of-domain source (The GenAI Divide) is handled. Applied 2026-07-27 for bp1-a1 and bp1-a3; the optional BP-06/08/09 edges remain unapplied.
Cautions and gaps¶
Author interest in the proposed remedy. The prescriptive section recommends the Agentic Automation Canvas, a framework authored by the same author (arXiv:2602.15090). The empirical review rests on independent trials and validations and stands on its own; the AAC recommendation is the author's own instrument and should be treated as one option for structured planning, not the only one. Any new practice drawn from this document should state the planning principle, not endorse a specific canvas.
Domain of the evidence. The primary evidence is software engineering copilots, ambient clinical scribes, and clinical decision support, not basic research workflows. The mechanisms (calibration error, verification burden, measurement mismatch, heterogeneity) are generic and transfer by argument; the specific figures should be cited as external field evidence, not as measurements of scientific practice. Clinical decision support and documentation are the closest touch to scientific and medical work.
Type of document. This is a narrative review, not new primary research, and an arXiv preprint that has not been peer reviewed. Its contribution is the aggregation and framing of others' controlled evidence.
What it does not claim. The document does not argue against agentic AI; it documents miscalibration and heterogeneity and states that "real gains exist in specific contexts and for specific user populations." Reading it as a case for non-adoption would misrepresent it.