0 cumulative citations
View corpus contextA deployment-grounded benchmark shows Knowtex’s clinical models had 97.99% of generated documentation tokens accepted into signed notes across over one million encounters and 13 specialties, but key companion statistics (notably completeness and signed-note rates) are withheld, limiting interpretation of the figure as true clinician effort reduction.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.
Summary
Main Finding
KnowBench introduces Effort Reduction (ER) — the proportion of a system-generated clinical artifact that the responsible clinician accepts under review — as a unified, deployment-grounded benchmark for clinical AI. Applied to Knowtex’s documentation system over a >6-month production window with >1,000,000 signed encounters across 13 specialties, the documentation-instantiation headline is: aggregate, content-weighted ER = 97.99% (strict variant counting formatting edits = 95.7%). The paper defines the metric, a six-item reporting protocol for auditable ER claims, instantiates ER across administrative tasks, and reports the documentation results while withholding some companion statistics (notably completeness and signed-note rate).
Key Points
- ER definition: ER(g,f) = |R(g,f)| / |g| where g = generated draft, f = clinician-finalized artifact, R(g,f) = retained content of g in f. Units depend on task (tokens, codes, order fields, recommendations).
- Primary operationalization for documentation: token-level retention computed via longest-common-subsequence alignment after formatting normalization.
- Companion ER variants:
- ER-text (retention, primary, cheap, computable from stored artifacts).
- ER-time (instrumented review time, used to validate ER-text and detect rubber-stamping).
- ER-cognitive (review/verification load, not directly instrumented).
- KnowBench protocol checklist (minimum for auditable claims): (1) measurement window & encounter count; (2) presentation rate & finalization rate; (3) completeness measure & floor; (4) formatting-edit handling; (5) unit & alignment method; (6) encounter-level distributional stats.
- Reported in this preprint: items (1), (4), (5) and presentation component of (2) — every generated draft in the window was presented and delivered to the EHR. Withheld: completeness measure, signed-note (finalization) rate, and encounter-level distributions.
- Documentation results (reported):
- Signed encounters: >1,000,000
- Measurement window: >6 contiguous months (ending H1 2026)
- Specialties: 13
- Aggregate ER (content-weighted): 97.99%
- Strict-variant ER (count formatting edits against retention): ≥95.7%
- Formatting-edit rate: 2.3% of tokens
- Per-specialty ER range: 96.8% (primary care, psychology) to 98.9% (nephrology)
- System context: Knowtex’s proprietary clinical foundation models with a closed feedback architecture (generation harness, per-clinician customization learned from signed notes, longitudinal conditioning on prior attested content) and continuous safety auditing.
- Important limitations stressed by authors:
- ER ≠ correctness: clinicians can sign erroneous content; ongoing factual-consistency audits are practiced but reported separately.
- ER can be inflated by carried-forward/templated content and by rubber-stamp review; completeness and ER-time are needed to address these confounds.
- Selected population: results reflect clinicians who adopted and continued using the system.
- Some companion statistics necessary for full auditability are withheld in this release.
Data & Methods
- Metric formalization:
- Unit: instantiation-specific (tokens for notes/summaries, codes for coding, fields for orders, recommendations for decision support).
- Alignment: token-level longest-common-subsequence after formatting normalization for documentation instantiation.
- Aggregation: content-weighted (so longer artifacts contribute proportionally).
- Validation instruments:
- ER-time (interaction-level instrumentation) recommended to validate that retention maps to actual time savings and to detect rubber-stamp behavior (high retention + near-zero review time).
- ER-cognitive discussed as aspirational.
- Protocol enforcement:
- Denominator discipline: encounters included only if a draft was generated, presented, reviewed, and signed (for the protocol in general). In this release, every generated draft was presented and delivered; the signed-note rate over delivered drafts is not reported.
- Formatting edits handled separately and excluded from primary ER figure; strict variant includes them.
- Companion statistics required (completeness, finalization rate, encounter-level distributions) to make ER claims auditable and comparable.
- Empirical dataset:
- Source: multiple anonymized provider organizations using Knowtex production deployment.
- Scale: >1,000,000 signed encounters across 13 specialties over >6 months.
- Architecture: closed-loop production system with per-clinician customization and longitudinal conditioning; corrections are fed back as training signal.
- Safety & audit:
- Knowtex runs a continuous factual-consistency audit program; results not fully reported here.
- Authors explicitly note the potential for automation bias and the need for independent safety demonstrations.
Implications for AI Economics
- Standardizing value measurement
- ER provides a concrete, deployment-grounded scalar tied to clinician acceptance that can reduce information asymmetry between vendors and purchasers. A standardized, auditable ER (with required companions) would allow purchasers to compare products on an operationally relevant metric rather than proxies (similarity scores, expert rubrics).
- However, withheld companions (completeness, signed-note rate, encounter distributions, ER-time) materially limit current auditability. For economic decision-making, buyers must require full protocol reporting to translate ER into financial benefit estimates.
- Mapping ER to productivity and dollars
- If validated (ER-text covaries with ER-time), ER can be used to estimate labor hours saved per encounter and thus compute ROI, payback periods, and per-encounter cost reductions — inputs needed for procurement, reimbursement negotiation, and CAPEX/OPEX modelling.
- Absent ER-time/ER-cognitive validation and completeness measures, mapping high ER to monetary value is risky: high token retention could reflect retained template text with little substantive time saved, carry-forward that shifts rather than saves effort, or signed-but-incorrect content that imposes downstream monitoring costs.
- Market structure, pricing, and contracting
- Closed-loop architectures with per-clinician customization (as in Knowtex) can yield higher ER and better realized productivity but increase switching costs and vendor lock-in, affecting competition and pricing power.
- Benchmarked ER (if standardized industry-wide) could support performance-based contracts (e.g., SLAs tied to ER and validated time-savings) or outcome-based procurement that ties fees to demonstrated effort reduction and safety metrics.
- Labor demand and task reallocation
- High ER (if matched by time-savings) implies reduced marginal document preparation labor, with implications for scribes, medical assistants, and billing coders. Economic outcomes could include:
- Reduced demand for some clerical roles, reallocation to higher-value tasks (care coordination, patient outreach), or contraction of workforce hours.
- Changes in wage dynamics and bargaining for clinicians as productivity per clinician rises (depending on how gains are captured by firms, clinicians, or payers).
- Complementarity effects: models that succeed at reducing documentation burden could increase clinician capacity to see more patients, potentially increasing throughput and revenue per clinician — contingent on non-documentation bottlenecks (scheduling, physical capacity, payer limits).
- Regulatory, liability and externality considerations
- Because ER measures acceptance (not correctness), insurers, regulators, and purchasers will require paired factual-consistency and safety audits before inferring economic value. Misaligned incentives or automation bias could externalize costs (medical errors, malpractice exposure) that negate claimed productivity gains.
- Public payers and regulators may demand standardized reporting (full KnowBench checklist) before approving broad reimbursements or incentives tied to AI-driven productivity.
- Macroeconomic measurement
- Adoption of ER-style metrics could change how healthcare productivity is measured in official statistics: a content-weighted effort-reduction metric tied to attestation events offers a more direct measure of automation’s impact on service-sector labor productivity than time-series EHR logs alone.
- Economists should be cautious: ER-based productivity improvements must be validated against realized clinician time, patient outcomes, and throughput to avoid overstating gains.
- Research and procurement recommendations
- Require full KnowBench reporting (all six checklist items) when using ER for procurement or economic evaluation.
- Insist on ER-time validation studies to translate ER-text into time and cost savings.
- Demand independent factual-consistency and safety audits to accompany ER claims.
- Model vendor lock-in effects due to per-clinician customization when assessing long-run costs and competition.
- Incorporate potential reallocation effects (task shifting, throughput changes) into benefit-cost analyses rather than assuming direct headcount reductions.
Summary takeaway: KnowBench and ER offer a promising, deployment-aligned way to quantify what clinical AI systems actually replace in clinicians’ workflows. For AI economics — procurement, pricing, ROI, labor impacts, and productivity accounting — ER could become a powerful instrument, but only if the full protocol (especially completeness, finalization rates, and ER-time validation) is reported and independently audited so acceptance maps reliably to real time- and outcome-valued savings.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| KnowBench defines Effort Reduction (ER) as the proportion of system-generated clinical work product accepted by the responsible clinician under expert review. Organizational Efficiency | positive | Reduction in clinician administrative effort through acceptance of AI-generated clinical work product |
Reading fidelity
high
Study strength
high
|
not reported
|
| Across more than one million signed encounters over a production window exceeding six months and spanning thirteen medical specialties, Knowtex's clinical AI system achieved an aggregate documentation ER of 97.99%. Organizational Efficiency | positive | Proportion of generated clinical documentation accepted without content edits and signed into the medical record |
Reading fidelity
high
Study strength
medium
|
n=1000000
97.99% aggregate ER
|
| Documentation ER varied across specialties from 96.8% to 98.9%. Organizational Efficiency | mixed | Specialty-specific acceptance rate of AI-generated documentation |
Reading fidelity
high
Study strength
medium
|
n=1000000
96.8–98.9% per-specialty ER
|
| When formatting-only edits are counted against retention, the strict-variant ER is at least 95.7%. Organizational Efficiency | positive | Documentation content retention under a stricter edit-counting rule |
Reading fidelity
high
Study strength
medium
|
n=1000000
95.7% strict-variant ER lower bound
|
| Formatting-only edits affected 2.3% of generated tokens in the measured documentation deployment. Error Rate | negative | Rate of generated documentation tokens requiring formatting edits |
Reading fidelity
high
Study strength
medium
|
n=1000000
2.3% of generated tokens
|
| Every draft generated during the measurement window was presented to the reviewing clinician and delivered to the EHR, rather than being selectively surfaced as a high-confidence subset. Organizational Efficiency | positive | Coverage of generated drafts entering clinician review |
Reading fidelity
high
Study strength
medium
|
n=1000000
100% of generated drafts presented
|
| The reported 97.99% ER figure is not sufficient by itself to establish correctness or safety, because clinician acceptance can include signing erroneous notes. Ai Safety And Ethics | mixed | Relationship between clinician acceptance of AI-generated notes and factual correctness or safety |
Reading fidelity
high
Study strength
high
|
not reported
|
| The preprint's ER result is only partially auditable because the finalization rate, completeness measure, and encounter-level distributional statistics are withheld. Governance And Regulation | negative | Auditability and interpretability of the reported ER benchmark result |
Reading fidelity
high
Study strength
high
|
not reported
|
| The reported 97.99% ER is a property of Knowtex's proprietary model family and closed-loop architecture, not of base models operating without those layers. Organizational Efficiency | mixed | Deployment-level documentation effort reduction attributable to the complete AI system architecture |
Reading fidelity
high
Study strength
high
|
n=1000000
97.99% ER for the closed-loop system
|