The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Independent audit finds roughly one in three AI-generated clinical notes contains at least one verified error, concentrated in allergies, medications and invented details; measured failure rates swing widely (≈25% to nearly all notes) depending on the reviewer standard and verification instrument.

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris · August 31, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Sebastian Fox unresolved corpus identity
  2. Luke Markham unresolved corpus identity
  3. Ryan Lail unresolved corpus identity
  4. Michael Karotsieris unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Sebastian Fox provider ID
  2. L. Markham provider ID
  3. Ryan Lail provider ID
  4. Michael Karotsieris provider ID
An adversarial, reproducible audit of three commercial ambient clinical scribes on 142 shared consultations found that 31.3% of notes contained at least one verified failure under the paper's strict standard, and that the measured error rate varies dramatically with the review instrument and standard.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.

Summary

Main Finding

A reproducible, adversarially verified audit of three commercial ambient AI scribes on the same 142 consultations (565 notes) finds that about one note in three contains at least one verified failure: 31.3% [27.0, 35.6]. The measured failure rate is highly sensitive to the audit instrument: reviewer instruction and choice of model family change verified rates dramatically (examples below). The authors publish all 618 verified findings, evidence quotes, prompts, model versions and the full pipeline so the census can be rerun under the same standard.

Key Points

  • Corpus and scope

    • 142 consultations → 565 notes across three commercial scribes (Scribes A–C).
    • Sources: 57 recorded UK primary-care (PriMock57), 45 US ambulatory (ACI-Bench), 40 authored scenarios (30 trap-seeded, 10 trap-blind).
    • Notes were generated without supplying patient records, demographics or encounter dates.
  • Discovery → Verification pipeline

    • 12 discovery passes (11 targeted + 1 open) produced 13,678 candidate errors; 5,898 cleared an importance filter.
    • Adversarial verification panel: two LLMs from different families (anthropic/claude-opus-5 = “harsher”, openai/gpt-5.5 = “gentler”) instructed to refute any candidate they could defensibly refute.
    • A third model (openai/gpt-5.4) served as a tiebreaker. A finding counted if both skeptics kept it or if tiebreak upheld it.
    • 618 findings survived verification (≈10.5% of candidates). Grouping suggests ≈265 distinct underlying errors (deduplication factor ~1.2–2.3).
  • Headline rates and error mix

    • Notes with ≥1 verified failure: 31.3% [27.0, 35.6].
    • If excluding failure classes a patient record would likely prefill (invented identities, invented dates), pooled rate = 24.8% [20.8, 29.0].
    • Common failure modes: wrong or missing allergy/medication information, invented patient identity/dates, and rewriting history as a physical exam on telephone consultations.
    • Authors identify a three-tier taxonomy aligned with published scribe-error taxonomies, plus one additional failure mode: recording a treatment as delivered after the clinician explicitly retracts it.
  • Human adjudication checks

    • Physician author blind-adjudicated 30 items (21 verified, 9 refused); upheld 20/21 verified findings → 95.2% precision [77.3, 99.2]. All 9 panel-refused candidates in that set were judged genuine failures by the physician (i.e., the panel had excluded them under its strict standard).
    • Independent clinician (not an author) adjudicated 16 items (12 verified, 4 refused); upheld all 12 verified → 100% [75.8, 100] and judged all 4 refused candidates genuine.
    • These results indicate the panel’s strict standard excludes some genuine errors (i.e., instrument decisions matter).
  • Instrument sensitivity (measured on a re-reviewed subset)

    • On a stratified fifth of candidates re-reviewed with everything else unchanged but a lenient reviewer instruction, the strict instruction verified 9.3% of those candidates while the lenient instruction verified 79.0%.
    • Under the same strict instruction, the gentler model family flagged 54.8% of sampled notes vs the harsher family’s 27.8% — roughly double. Adding the gentler family as a second opinion altered the headline by ≈1 percentage point.
    • Depending on the audit standard applied, between 28% and 97% of sampled notes can be considered to carry a verified failure.
  • Comparison to prior audits

    • Published audits report widely differing error mixes (omission dominant at 54–86% of errors). This paper shows instrument differences of the measured magnitude can produce those disagreements; when counting notes as this paper does, omission shares are similar to some prior work (e.g., one study reports 18% omission over notes vs this paper’s 15.4%).
  • Transparency and reproducibility

    • All 618 verified findings, transcript-side evidence, every prompt and model version, and the re-runnable pipeline are released so other researchers or buyers can repeat the census under the same published standard.

Data & Methods

  • Corpus: 142 consultations; notes: 282 (Scribe A), 141 (Scribe B), 142 (Scribe C), total 565 notes.
  • Discovery stage: 12 passes (11 targeted hunts for known failure modes + 1 open pass), run on anthropic/claude-opus-5; produced 13,678 candidate error annotations; 5,898 passed importance filter.
  • Verification stage: adversarial "skeptics" (claude-opus-5 and gpt-5.5), each shown full note + full transcript and instructed to refute candidates whenever defensible. A third model (gpt-5.4 with high reasoning effort) served as tiebreaker. Deterministic seeds used for reproducibility; exact settings and manifests released.
  • Outcome rules: a candidate counts as a verified finding if both skeptics keep it or if the tiebreaker upholds it; unparseable/failed responses count as refutations.
  • Human validation: two blinded human clinicians independently adjudicated disjoint random samples of verified findings + some high-importance refused candidates to assess instrument precision and the panel’s refusal decisions.
  • Grouping/deduplication: findings are not merged across passes; model-assisted grouping and hand checks estimate 265 distinct errors among 618 findings.
  • Release: full dataset of findings, evidence quotes, prompts, model versions, and runnable pipeline.

Implications for AI Economics

  • Measurement is a market and regulatory input
    • Reported error rates for AI scribes are not intrinsic product properties but joint outcomes of product behavior and the audit instrument (discovery heuristics, reviewer instruction, model family). Procurement, regulation, and reimbursement decisions that use headline error rates need to require standardized, transparent instruments or run reproducible third-party audits.
  • Competitive signaling and perverse incentives
    • Vendors may benefit from publishing favorable audit instruments or calibrating systems to perform well on specific detectors. Absent transparent, shared standards, vendors can game audits (optimize for detector sensitivity, not clinical utility), complicating buyer comparisons.
  • Liability, pricing and insurance
    • Variation in measured failure rates changes expected downstream liability and thus insurance pricing and legal risk assessments. Purchasers and payers should factor audit-instrument uncertainty into contract terms, warranty clauses, indemnities and due-diligence processes.
  • Product design and integration economics
    • Many errors (invented identities/dates, some omissions) arise or would be mitigated by integrating patient records and structured context. That creates economic value for tighter EHR integration and potentially lock-in: integrated solutions can reduce certain failure classes and therefore reduce oversight costs.
  • Cost of oversight and clinical labor
    • The census underscores that clinician sign-off matters (clinicians authoritatively control notes at point-of-signature). But verifying and correcting AI-suggested notes imposes monitoring costs; economics of adoption depend on how often clinicians must intervene and the time-cost per intervention. Buyers must evaluate total cost of ownership (AI + associated clinician QA) rather than vendor-reported headline error rates alone.
  • Role for reproducible third-party audits as public goods
    • The authors’ open, re-runnable instrument is a model for standardized third-party assessments. Publicly available, reproducible audits reduce information asymmetries between vendors and buyers, improve market transparency, and help regulators set evidence-based standards.
  • Policy and regulation
    • Regulators and standard-setting bodies should require disclosure of audit instruments and promote benchmarked evaluation protocols. Regulatory decisions (approvals, sandbox outcomes, labelling) that rest on opaque internal audits risk being misled by instrument-dependent numbers.
  • Investment and product strategy implications
    • Investors and procurement officers should demand audits with transparent, reproducible instruments and take care when comparing vendor claims. Vendors that support open benchmarking and clear auditability may enjoy trust premiums; conversely, those optimizing for narrowly defined detectors may face reputational and regulatory risk.

Bottom line: error rates for deployed AI scribes are meaningful only when reported alongside the exact, reproducible audit instrument that produced them. For economic decisions—procurement, pricing, liability, and regulation—buyers and policymakers must treat the instrument as part of the measurement and demand transparency and standardized benchmarks.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides a carefully constructed, reproducible audit with adversarial LLM verification and blinded clinician spot-checks, producing a sizeable set of verified findings (618) from 565 notes; however, the evidence applies to three specific deployed products run without EHR context, with dated/unnamed product snapshots and non-live captures, limiting external validity beyond the studied configurations. Methods Rigorhigh — The study uses over-inclusive discovery passes, adversarial verification by two different LLM families plus a tiebreaker, stratified corpus construction (real recorded encounters plus authored scenarios), systematic importance filtering, and blinded human adjudication; it also tests the sensitivity of the measured error rate to reviewer instruction and model family and releases the entire instrument and evidence for replication. Limitations include lack of product versioning, non-live deployment (no integrated EHR context), and no preregistration mentioned. Sample142 consultations (57 UK primary-care PriMock57, 45 US ambulatory ACI-Bench, 40 authored scenarios), producing 565 notes from three commercial ambient scribes (282 notes Scribe A, 141 Scribe B, 142 Scribe C). Discovery: 12 passes produced 13,678 candidate errors, 5,898 passed an importance filter. Verification: adversarial panel of two LLM families (anthropic/claude-opus-5 and openai/gpt-5.5) with openai/gpt-5.4 as tiebreaker verified 618 findings. Human blinded adjudication: a physician-author and an independent clinician reviewed disjoint random samples. Notes were generated by replayed audio for some products and transcript API for one; no patient records/demographics were provided to products. Captures occurred June–August 2026; product versions were not available. Themeshuman_ai_collab governance GeneralizabilityOnly three commercial scribe products evaluated — results may not generalize across other vendors or newer versions, Notes generated without EHR/patient-record context; integrated deployments that prefill patient data could change error patterns, Corpus mixes recorded encounters and authored scenarios; authored trap scenarios may not reflect real-world frequency of errors, Product snapshots are dated and unversioned (June–August 2026); results may not hold for later model updates, Sample sizes by stratum are modest; some strata (e.g., trap-blind authored scenarios) are small, Findings measure failures under a specified verification instrument and reviewer standard; different instruments produce substantially different headline rates

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Across three commercial ambient AI scribe products, 31.3% of notes contained at least one verified failure. Error Rate negative Presence of at least one verified scribe failure per clinical note
Reading fidelity high
Study strength medium
n=565
31.3% [27.0, 35.6]
0.18
The verified failures concentrated in allergy and medication information, invented patient identity, and physical examinations documented for telephone consultations that could not have included an examination. Error Rate negative Types and distribution of documentation failures
Reading fidelity high
Study strength medium
n=618
0.18
When no patient record was provided, the rate of notes containing at least one verified failure was 24.8% after setting aside invented identities and invented dates that a record would have prefilled. Error Rate negative Rate of notes containing at least one verified failure under a record-adjusted definition
Reading fidelity high
Study strength medium
n=565
24.8% [20.8, 29.0]
0.18
The verification panel retained 618 of 5,898 candidate errors, corresponding to a verified-candidate rate of 10.48%. Error Rate negative Proportion of candidate errors verified by the review panel
Reading fidelity high
Study strength medium
n=5898
10.48% [8.9, 12.1]
0.18
Changing only the review instruction from strict to lenient increased the share of candidates verified from 9.3% to 79.0%. Error Rate positive Share of candidate errors verified under different review standards
Reading fidelity high
Study strength high
n=1295
9.3% to 79.0%
0.3
At the same strict instruction, the gentler reviewing model family flagged 54.8% of sampled notes compared with 27.8% for the harsher family. Error Rate positive Share of sampled notes flagged as containing a verified failure
Reading fidelity high
Study strength high
54.8% versus 27.8%
0.3
Adding the gentler model family as a second opinion with a tiebreak increased the count by only one percentage point relative to the harsher family alone. Error Rate positive Change in the proportion of notes counted as containing a verified failure
Reading fidelity high
Study strength high
one percentage point
0.3
Depending on the review standard applied, between 28% and 97% of sampled notes carried a verified failure. Error Rate mixed Rate of sampled notes containing a verified failure under alternative counting instruments
Reading fidelity high
Study strength high
28% to 97%
0.3
A physician author upheld 20 of 21 sampled verified findings, while an independent clinician upheld all 12 of 12 sampled verified findings. Ai Safety And Ethics positive Human agreement that AI-panel-verified findings were genuine failures
Reading fidelity high
Study strength low
n=33
20/21 (95.2% [77.3, 99.2]) and 12/12 ([75.8, 100])
0.09
The omission share in this audit was 23.1%, substantially lower than the 54%–86% omission share reported across the cited published audits. Error Rate mixed Share of identified failures classified as omissions
Reading fidelity high
Study strength low
23.1% versus 54%–86%
0.09
The trap-seeded authored scenarios produced more findings per note than the trap-blind authored scenarios, but the difference was not statistically distinguishable from zero at the reported sample sizes. Error Rate null_result Verified findings per note in authored scenarios
Reading fidelity high
Study strength medium
n=157
0.475 versus 0.297 findings per note
0.18

Notes