The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A deployment-grounded benchmark shows Knowtex’s clinical models had 97.99% of generated documentation tokens accepted into signed notes across over one million encounters and 13 specialties, but key companion statistics (notably completeness and signed-note rates) are withheld, limiting interpretation of the figure as true clinician effort reduction.

KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
Jocelyn Kang, Caroline Zhang · September 14, 2026
arxiv descriptive medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jocelyn Kang unresolved corpus identity
  2. Caroline Zhang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jocelyn Kang unresolved corpus identity
  2. Caroline Zhang unresolved corpus identity
KnowBench defines Effort Reduction (ER) as the proportion of system-generated clinical content accepted by the reviewing clinician and reports a documentation-instantiation ER of 97.99% across >1,000,000 signed encounters using Knowtex's proprietary models, while withholding some companion statistics required by the protocol.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.

Summary

Main Finding

KnowBench introduces Effort Reduction (ER) — the proportion of a system-generated clinical artifact that the responsible clinician accepts under review — as a unified, deployment-grounded benchmark for clinical AI. Applied to Knowtex’s documentation system over a >6-month production window with >1,000,000 signed encounters across 13 specialties, the documentation-instantiation headline is: aggregate, content-weighted ER = 97.99% (strict variant counting formatting edits = 95.7%). The paper defines the metric, a six-item reporting protocol for auditable ER claims, instantiates ER across administrative tasks, and reports the documentation results while withholding some companion statistics (notably completeness and signed-note rate).

Key Points

  • ER definition: ER(g,f) = |R(g,f)| / |g| where g = generated draft, f = clinician-finalized artifact, R(g,f) = retained content of g in f. Units depend on task (tokens, codes, order fields, recommendations).
  • Primary operationalization for documentation: token-level retention computed via longest-common-subsequence alignment after formatting normalization.
  • Companion ER variants:
    • ER-text (retention, primary, cheap, computable from stored artifacts).
    • ER-time (instrumented review time, used to validate ER-text and detect rubber-stamping).
    • ER-cognitive (review/verification load, not directly instrumented).
  • KnowBench protocol checklist (minimum for auditable claims): (1) measurement window & encounter count; (2) presentation rate & finalization rate; (3) completeness measure & floor; (4) formatting-edit handling; (5) unit & alignment method; (6) encounter-level distributional stats.
  • Reported in this preprint: items (1), (4), (5) and presentation component of (2) — every generated draft in the window was presented and delivered to the EHR. Withheld: completeness measure, signed-note (finalization) rate, and encounter-level distributions.
  • Documentation results (reported):
    • Signed encounters: >1,000,000
    • Measurement window: >6 contiguous months (ending H1 2026)
    • Specialties: 13
    • Aggregate ER (content-weighted): 97.99%
    • Strict-variant ER (count formatting edits against retention): ≥95.7%
    • Formatting-edit rate: 2.3% of tokens
    • Per-specialty ER range: 96.8% (primary care, psychology) to 98.9% (nephrology)
  • System context: Knowtex’s proprietary clinical foundation models with a closed feedback architecture (generation harness, per-clinician customization learned from signed notes, longitudinal conditioning on prior attested content) and continuous safety auditing.
  • Important limitations stressed by authors:
    • ER ≠ correctness: clinicians can sign erroneous content; ongoing factual-consistency audits are practiced but reported separately.
    • ER can be inflated by carried-forward/templated content and by rubber-stamp review; completeness and ER-time are needed to address these confounds.
    • Selected population: results reflect clinicians who adopted and continued using the system.
    • Some companion statistics necessary for full auditability are withheld in this release.

Data & Methods

  • Metric formalization:
    • Unit: instantiation-specific (tokens for notes/summaries, codes for coding, fields for orders, recommendations for decision support).
    • Alignment: token-level longest-common-subsequence after formatting normalization for documentation instantiation.
    • Aggregation: content-weighted (so longer artifacts contribute proportionally).
  • Validation instruments:
    • ER-time (interaction-level instrumentation) recommended to validate that retention maps to actual time savings and to detect rubber-stamp behavior (high retention + near-zero review time).
    • ER-cognitive discussed as aspirational.
  • Protocol enforcement:
    • Denominator discipline: encounters included only if a draft was generated, presented, reviewed, and signed (for the protocol in general). In this release, every generated draft was presented and delivered; the signed-note rate over delivered drafts is not reported.
    • Formatting edits handled separately and excluded from primary ER figure; strict variant includes them.
    • Companion statistics required (completeness, finalization rate, encounter-level distributions) to make ER claims auditable and comparable.
  • Empirical dataset:
    • Source: multiple anonymized provider organizations using Knowtex production deployment.
    • Scale: >1,000,000 signed encounters across 13 specialties over >6 months.
    • Architecture: closed-loop production system with per-clinician customization and longitudinal conditioning; corrections are fed back as training signal.
  • Safety & audit:
    • Knowtex runs a continuous factual-consistency audit program; results not fully reported here.
    • Authors explicitly note the potential for automation bias and the need for independent safety demonstrations.

Implications for AI Economics

  • Standardizing value measurement
    • ER provides a concrete, deployment-grounded scalar tied to clinician acceptance that can reduce information asymmetry between vendors and purchasers. A standardized, auditable ER (with required companions) would allow purchasers to compare products on an operationally relevant metric rather than proxies (similarity scores, expert rubrics).
    • However, withheld companions (completeness, signed-note rate, encounter distributions, ER-time) materially limit current auditability. For economic decision-making, buyers must require full protocol reporting to translate ER into financial benefit estimates.
  • Mapping ER to productivity and dollars
    • If validated (ER-text covaries with ER-time), ER can be used to estimate labor hours saved per encounter and thus compute ROI, payback periods, and per-encounter cost reductions — inputs needed for procurement, reimbursement negotiation, and CAPEX/OPEX modelling.
    • Absent ER-time/ER-cognitive validation and completeness measures, mapping high ER to monetary value is risky: high token retention could reflect retained template text with little substantive time saved, carry-forward that shifts rather than saves effort, or signed-but-incorrect content that imposes downstream monitoring costs.
  • Market structure, pricing, and contracting
    • Closed-loop architectures with per-clinician customization (as in Knowtex) can yield higher ER and better realized productivity but increase switching costs and vendor lock-in, affecting competition and pricing power.
    • Benchmarked ER (if standardized industry-wide) could support performance-based contracts (e.g., SLAs tied to ER and validated time-savings) or outcome-based procurement that ties fees to demonstrated effort reduction and safety metrics.
  • Labor demand and task reallocation
    • High ER (if matched by time-savings) implies reduced marginal document preparation labor, with implications for scribes, medical assistants, and billing coders. Economic outcomes could include:
    • Reduced demand for some clerical roles, reallocation to higher-value tasks (care coordination, patient outreach), or contraction of workforce hours.
    • Changes in wage dynamics and bargaining for clinicians as productivity per clinician rises (depending on how gains are captured by firms, clinicians, or payers).
    • Complementarity effects: models that succeed at reducing documentation burden could increase clinician capacity to see more patients, potentially increasing throughput and revenue per clinician — contingent on non-documentation bottlenecks (scheduling, physical capacity, payer limits).
  • Regulatory, liability and externality considerations
    • Because ER measures acceptance (not correctness), insurers, regulators, and purchasers will require paired factual-consistency and safety audits before inferring economic value. Misaligned incentives or automation bias could externalize costs (medical errors, malpractice exposure) that negate claimed productivity gains.
    • Public payers and regulators may demand standardized reporting (full KnowBench checklist) before approving broad reimbursements or incentives tied to AI-driven productivity.
  • Macroeconomic measurement
    • Adoption of ER-style metrics could change how healthcare productivity is measured in official statistics: a content-weighted effort-reduction metric tied to attestation events offers a more direct measure of automation’s impact on service-sector labor productivity than time-series EHR logs alone.
    • Economists should be cautious: ER-based productivity improvements must be validated against realized clinician time, patient outcomes, and throughput to avoid overstating gains.
  • Research and procurement recommendations
    • Require full KnowBench reporting (all six checklist items) when using ER for procurement or economic evaluation.
    • Insist on ER-time validation studies to translate ER-text into time and cost savings.
    • Demand independent factual-consistency and safety audits to accompany ER claims.
    • Model vendor lock-in effects due to per-clinician customization when assessing long-run costs and competition.
    • Incorporate potential reallocation effects (task shifting, throughput changes) into benefit-cost analyses rather than assuming direct headcount reductions.

Summary takeaway: KnowBench and ER offer a promising, deployment-aligned way to quantify what clinical AI systems actually replace in clinicians’ workflows. For AI economics — procurement, pricing, ROI, labor impacts, and productivity accounting — ER could become a powerful instrument, but only if the full protocol (especially completeness, finalization rates, and ER-time validation) is reported and independently audited so acceptance maps reliably to real time- and outcome-valued savings.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper reports very large-scale deployment measurements (>1M signed encounters across >6 months and 13 specialties) which give empirical weight to the claims, and it formalizes a clear, interoperable metric (ER) and protocol. However, several key companion statistics required by the protocol (completeness and signed-note-rate, encounter-level distributions) are withheld in the preprint, and important confounds (automation bias, carry-forward/ note-bloat, adopter selection) remain only partially addressed, limiting confidence in interpreting ER as true effort reduction. Methods Rigormedium — The authors provide a formal metric definition, alignment and tokenization choices, validation concepts (ER-time, factual audits), and a reporting checklist—indicative of strong measurement thinking. But the empirical release omits several protocol companions, relies on proprietary closed-loop models and internal audits, and does not publish the distributional or completeness data needed to rule out major confounds, reducing methodological transparency and reproducibility. SampleProprietary deployment of Knowtex fine-tuned clinical foundation models operating inside a closed feedback architecture; measurement window >6 contiguous months ending H1 2026; denominator >1,000,000 clinician-signed encounters across 13 medical specialties at multiple anonymized provider organizations; every generated draft in the window was presented to clinicians and delivered to the EHR; retention computed at token level (LCS alignment) with formatting edits reported separately; companion statistics (completeness, signed-note-rate, encounter-level distributions) withheld from the preprint. Themesproductivity human_ai_collab GeneralizabilityMeasurement reflects a proprietary closed-loop architecture with per-clinician customization, so results may not generalize to base LLMs or other deployment models., Population limited to clinicians who adopted and continued using the system (selection/retention bias)., Thirteen specialties reported but not necessarily representative of all clinical settings or international health systems., Withheld completeness and signed-note-rate prevent assessing whether high ER reflects substantive automation or templated/ carried-forward content (note-bloat)., Metric depends on specific workflow/EHR integration and review incentives; results may not transfer to different workflows or regulatory contexts.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
KnowBench defines Effort Reduction (ER) as the proportion of system-generated clinical work product accepted by the responsible clinician under expert review. Organizational Efficiency positive Reduction in clinician administrative effort through acceptance of AI-generated clinical work product
Reading fidelity high
Study strength high
not reported
0.3
Across more than one million signed encounters over a production window exceeding six months and spanning thirteen medical specialties, Knowtex's clinical AI system achieved an aggregate documentation ER of 97.99%. Organizational Efficiency positive Proportion of generated clinical documentation accepted without content edits and signed into the medical record
Reading fidelity high
Study strength medium
n=1000000
97.99% aggregate ER
0.18
Documentation ER varied across specialties from 96.8% to 98.9%. Organizational Efficiency mixed Specialty-specific acceptance rate of AI-generated documentation
Reading fidelity high
Study strength medium
n=1000000
96.8–98.9% per-specialty ER
0.18
When formatting-only edits are counted against retention, the strict-variant ER is at least 95.7%. Organizational Efficiency positive Documentation content retention under a stricter edit-counting rule
Reading fidelity high
Study strength medium
n=1000000
95.7% strict-variant ER lower bound
0.18
Formatting-only edits affected 2.3% of generated tokens in the measured documentation deployment. Error Rate negative Rate of generated documentation tokens requiring formatting edits
Reading fidelity high
Study strength medium
n=1000000
2.3% of generated tokens
0.18
Every draft generated during the measurement window was presented to the reviewing clinician and delivered to the EHR, rather than being selectively surfaced as a high-confidence subset. Organizational Efficiency positive Coverage of generated drafts entering clinician review
Reading fidelity high
Study strength medium
n=1000000
100% of generated drafts presented
0.18
The reported 97.99% ER figure is not sufficient by itself to establish correctness or safety, because clinician acceptance can include signing erroneous notes. Ai Safety And Ethics mixed Relationship between clinician acceptance of AI-generated notes and factual correctness or safety
Reading fidelity high
Study strength high
not reported
0.3
The preprint's ER result is only partially auditable because the finalization rate, completeness measure, and encounter-level distributional statistics are withheld. Governance And Regulation negative Auditability and interpretability of the reported ER benchmark result
Reading fidelity high
Study strength high
not reported
0.3
The reported 97.99% ER is a property of Knowtex's proprietary model family and closed-loop architecture, not of base models operating without those layers. Organizational Efficiency mixed Deployment-level documentation effort reduction attributable to the complete AI system architecture
Reading fidelity high
Study strength high
n=1000000
97.99% ER for the closed-loop system
0.3

Notes