Independent audit finds roughly one in three AI-generated clinical notes contains at least one verified error, concentrated in allergies, medications and invented details; measured failure rates swing widely (≈25% to nearly all notes) depending on the reviewer standard and verification instrument.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.
Summary
Main Finding
A reproducible, adversarially verified audit of three commercial ambient AI scribes on the same 142 consultations (565 notes) finds that about one note in three contains at least one verified failure: 31.3% [27.0, 35.6]. The measured failure rate is highly sensitive to the audit instrument: reviewer instruction and choice of model family change verified rates dramatically (examples below). The authors publish all 618 verified findings, evidence quotes, prompts, model versions and the full pipeline so the census can be rerun under the same standard.
Key Points
-
Corpus and scope
- 142 consultations → 565 notes across three commercial scribes (Scribes A–C).
- Sources: 57 recorded UK primary-care (PriMock57), 45 US ambulatory (ACI-Bench), 40 authored scenarios (30 trap-seeded, 10 trap-blind).
- Notes were generated without supplying patient records, demographics or encounter dates.
-
Discovery → Verification pipeline
- 12 discovery passes (11 targeted + 1 open) produced 13,678 candidate errors; 5,898 cleared an importance filter.
- Adversarial verification panel: two LLMs from different families (anthropic/claude-opus-5 = “harsher”, openai/gpt-5.5 = “gentler”) instructed to refute any candidate they could defensibly refute.
- A third model (openai/gpt-5.4) served as a tiebreaker. A finding counted if both skeptics kept it or if tiebreak upheld it.
- 618 findings survived verification (≈10.5% of candidates). Grouping suggests ≈265 distinct underlying errors (deduplication factor ~1.2–2.3).
-
Headline rates and error mix
- Notes with ≥1 verified failure: 31.3% [27.0, 35.6].
- If excluding failure classes a patient record would likely prefill (invented identities, invented dates), pooled rate = 24.8% [20.8, 29.0].
- Common failure modes: wrong or missing allergy/medication information, invented patient identity/dates, and rewriting history as a physical exam on telephone consultations.
- Authors identify a three-tier taxonomy aligned with published scribe-error taxonomies, plus one additional failure mode: recording a treatment as delivered after the clinician explicitly retracts it.
-
Human adjudication checks
- Physician author blind-adjudicated 30 items (21 verified, 9 refused); upheld 20/21 verified findings → 95.2% precision [77.3, 99.2]. All 9 panel-refused candidates in that set were judged genuine failures by the physician (i.e., the panel had excluded them under its strict standard).
- Independent clinician (not an author) adjudicated 16 items (12 verified, 4 refused); upheld all 12 verified → 100% [75.8, 100] and judged all 4 refused candidates genuine.
- These results indicate the panel’s strict standard excludes some genuine errors (i.e., instrument decisions matter).
-
Instrument sensitivity (measured on a re-reviewed subset)
- On a stratified fifth of candidates re-reviewed with everything else unchanged but a lenient reviewer instruction, the strict instruction verified 9.3% of those candidates while the lenient instruction verified 79.0%.
- Under the same strict instruction, the gentler model family flagged 54.8% of sampled notes vs the harsher family’s 27.8% — roughly double. Adding the gentler family as a second opinion altered the headline by ≈1 percentage point.
- Depending on the audit standard applied, between 28% and 97% of sampled notes can be considered to carry a verified failure.
-
Comparison to prior audits
- Published audits report widely differing error mixes (omission dominant at 54–86% of errors). This paper shows instrument differences of the measured magnitude can produce those disagreements; when counting notes as this paper does, omission shares are similar to some prior work (e.g., one study reports 18% omission over notes vs this paper’s 15.4%).
-
Transparency and reproducibility
- All 618 verified findings, transcript-side evidence, every prompt and model version, and the re-runnable pipeline are released so other researchers or buyers can repeat the census under the same published standard.
Data & Methods
- Corpus: 142 consultations; notes: 282 (Scribe A), 141 (Scribe B), 142 (Scribe C), total 565 notes.
- Discovery stage: 12 passes (11 targeted hunts for known failure modes + 1 open pass), run on anthropic/claude-opus-5; produced 13,678 candidate error annotations; 5,898 passed importance filter.
- Verification stage: adversarial "skeptics" (claude-opus-5 and gpt-5.5), each shown full note + full transcript and instructed to refute candidates whenever defensible. A third model (gpt-5.4 with high reasoning effort) served as tiebreaker. Deterministic seeds used for reproducibility; exact settings and manifests released.
- Outcome rules: a candidate counts as a verified finding if both skeptics keep it or if the tiebreaker upholds it; unparseable/failed responses count as refutations.
- Human validation: two blinded human clinicians independently adjudicated disjoint random samples of verified findings + some high-importance refused candidates to assess instrument precision and the panel’s refusal decisions.
- Grouping/deduplication: findings are not merged across passes; model-assisted grouping and hand checks estimate 265 distinct errors among 618 findings.
- Release: full dataset of findings, evidence quotes, prompts, model versions, and runnable pipeline.
Implications for AI Economics
- Measurement is a market and regulatory input
- Reported error rates for AI scribes are not intrinsic product properties but joint outcomes of product behavior and the audit instrument (discovery heuristics, reviewer instruction, model family). Procurement, regulation, and reimbursement decisions that use headline error rates need to require standardized, transparent instruments or run reproducible third-party audits.
- Competitive signaling and perverse incentives
- Vendors may benefit from publishing favorable audit instruments or calibrating systems to perform well on specific detectors. Absent transparent, shared standards, vendors can game audits (optimize for detector sensitivity, not clinical utility), complicating buyer comparisons.
- Liability, pricing and insurance
- Variation in measured failure rates changes expected downstream liability and thus insurance pricing and legal risk assessments. Purchasers and payers should factor audit-instrument uncertainty into contract terms, warranty clauses, indemnities and due-diligence processes.
- Product design and integration economics
- Many errors (invented identities/dates, some omissions) arise or would be mitigated by integrating patient records and structured context. That creates economic value for tighter EHR integration and potentially lock-in: integrated solutions can reduce certain failure classes and therefore reduce oversight costs.
- Cost of oversight and clinical labor
- The census underscores that clinician sign-off matters (clinicians authoritatively control notes at point-of-signature). But verifying and correcting AI-suggested notes imposes monitoring costs; economics of adoption depend on how often clinicians must intervene and the time-cost per intervention. Buyers must evaluate total cost of ownership (AI + associated clinician QA) rather than vendor-reported headline error rates alone.
- Role for reproducible third-party audits as public goods
- The authors’ open, re-runnable instrument is a model for standardized third-party assessments. Publicly available, reproducible audits reduce information asymmetries between vendors and buyers, improve market transparency, and help regulators set evidence-based standards.
- Policy and regulation
- Regulators and standard-setting bodies should require disclosure of audit instruments and promote benchmarked evaluation protocols. Regulatory decisions (approvals, sandbox outcomes, labelling) that rest on opaque internal audits risk being misled by instrument-dependent numbers.
- Investment and product strategy implications
- Investors and procurement officers should demand audits with transparent, reproducible instruments and take care when comparing vendor claims. Vendors that support open benchmarking and clear auditability may enjoy trust premiums; conversely, those optimizing for narrowly defined detectors may face reputational and regulatory risk.
Bottom line: error rates for deployed AI scribes are meaningful only when reported alongside the exact, reproducible audit instrument that produced them. For economic decisions—procurement, pricing, liability, and regulation—buyers and policymakers must treat the instrument as part of the measurement and demand transparency and standardized benchmarks.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Across three commercial ambient AI scribe products, 31.3% of notes contained at least one verified failure. Error Rate | negative | Presence of at least one verified scribe failure per clinical note |
Reading fidelity
high
Study strength
medium
|
n=565
31.3% [27.0, 35.6]
|
| The verified failures concentrated in allergy and medication information, invented patient identity, and physical examinations documented for telephone consultations that could not have included an examination. Error Rate | negative | Types and distribution of documentation failures |
Reading fidelity
high
Study strength
medium
|
n=618
|
| When no patient record was provided, the rate of notes containing at least one verified failure was 24.8% after setting aside invented identities and invented dates that a record would have prefilled. Error Rate | negative | Rate of notes containing at least one verified failure under a record-adjusted definition |
Reading fidelity
high
Study strength
medium
|
n=565
24.8% [20.8, 29.0]
|
| The verification panel retained 618 of 5,898 candidate errors, corresponding to a verified-candidate rate of 10.48%. Error Rate | negative | Proportion of candidate errors verified by the review panel |
Reading fidelity
high
Study strength
medium
|
n=5898
10.48% [8.9, 12.1]
|
| Changing only the review instruction from strict to lenient increased the share of candidates verified from 9.3% to 79.0%. Error Rate | positive | Share of candidate errors verified under different review standards |
Reading fidelity
high
Study strength
high
|
n=1295
9.3% to 79.0%
|
| At the same strict instruction, the gentler reviewing model family flagged 54.8% of sampled notes compared with 27.8% for the harsher family. Error Rate | positive | Share of sampled notes flagged as containing a verified failure |
Reading fidelity
high
Study strength
high
|
54.8% versus 27.8%
|
| Adding the gentler model family as a second opinion with a tiebreak increased the count by only one percentage point relative to the harsher family alone. Error Rate | positive | Change in the proportion of notes counted as containing a verified failure |
Reading fidelity
high
Study strength
high
|
one percentage point
|
| Depending on the review standard applied, between 28% and 97% of sampled notes carried a verified failure. Error Rate | mixed | Rate of sampled notes containing a verified failure under alternative counting instruments |
Reading fidelity
high
Study strength
high
|
28% to 97%
|
| A physician author upheld 20 of 21 sampled verified findings, while an independent clinician upheld all 12 of 12 sampled verified findings. Ai Safety And Ethics | positive | Human agreement that AI-panel-verified findings were genuine failures |
Reading fidelity
high
Study strength
low
|
n=33
20/21 (95.2% [77.3, 99.2]) and 12/12 ([75.8, 100])
|
| The omission share in this audit was 23.1%, substantially lower than the 54%–86% omission share reported across the cited published audits. Error Rate | mixed | Share of identified failures classified as omissions |
Reading fidelity
high
Study strength
low
|
23.1% versus 54%–86%
|
| The trap-seeded authored scenarios produced more findings per note than the trap-blind authored scenarios, but the difference was not statistically distinguishable from zero at the reported sample sizes. Error Rate | null_result | Verified findings per note in authored scenarios |
Reading fidelity
high
Study strength
medium
|
n=157
0.475 versus 0.297 findings per note
|