The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

An explainable LLM system for oncology trial matching found clinician-recommended trials for the vast majority of retrospective cases and halved screening time, and in a six-month tumor-board pilot it surfaced many trial options missed by routine review, substantially expanding potential trial access.

Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment
Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh, Ryan Huu-Tuan Nguyen, Zifeng Wang, Jimeng Sun, Charalampos S. Floudas, James L. Gulley, Kamilia Moalem, Catarina Martins Maia, Amanda Nottke, Juan W. Valle, Melinda Bachini, Lourdes Rocha-Nussbaum, Kari Ramage, Nikita Curry, Megan Barnes, Mandy Mansaray, Darlene Gabeau, Craig E. Grossman, Heath Skinner, Michael Burczynski, NIH-TrialBench Consortium, Zhiyong Lu · September 01, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yin Fang unresolved corpus identity
  2. Qiao Jin unresolved corpus identity
  3. Shubo Tian unresolved corpus identity
  4. Lauren He unresolved corpus identity
  5. Maya Geer unresolved corpus identity
  6. Noor Naffakh unresolved corpus identity
  7. Ryan Huu-Tuan Nguyen unresolved corpus identity
  8. Zifeng Wang unresolved corpus identity
  9. Jimeng Sun unresolved corpus identity
  10. Charalampos S. Floudas unresolved corpus identity
  11. James L. Gulley unresolved corpus identity
  12. Kamilia Moalem unresolved corpus identity
  13. Catarina Martins Maia unresolved corpus identity
  14. Amanda Nottke unresolved corpus identity
  15. Juan W. Valle unresolved corpus identity
  16. Melinda Bachini unresolved corpus identity
  17. Lourdes Rocha-Nussbaum unresolved corpus identity
  18. Kari Ramage unresolved corpus identity
  19. Nikita Curry unresolved corpus identity
  20. Megan Barnes unresolved corpus identity
  21. Mandy Mansaray unresolved corpus identity
  22. Darlene Gabeau unresolved corpus identity
  23. Craig E. Grossman unresolved corpus identity
  24. Heath Skinner unresolved corpus identity
  25. Michael Burczynski unresolved corpus identity
  26. NIH-TrialBench Consortium unresolved corpus identity
  27. Zhiyong Lu unresolved corpus identity
TrialGPT 2.0, an LLM-based, explainable clinical trial matching system, retrieved clinician-recommended trials for ~91% of retrospective cases, reduced clinician screening time by ~55%, and in a six-month prospective tumor-board pilot contributed additional trial opportunities, increasing potential access by 90.9%.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world deployment. Rather than asking only whether a patient may qualify, the system also assesses which trials warrant further consideration given the patient's current clinical needs and local workflow priorities, and provides structured, inspectable explanations for expert review. Importantly, we evaluated TrialGPT 2.0 retrospectively and prospectively across multiple oncology-focused settings, spanning government, academic cancer-center, patient-advocacy, and NIH referral workflows. In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial in its top 10 recommendations for approximately 91% of cases while reducing clinician screening time by 55.0%. In a six-month prospective evaluation embedded in an active precision oncology tumor board, TrialGPT 2.0 contributed additional trial opportunities missed by the routine workflow, expanding patient access to clinical trial participation by 90.9%. To support scientific reproducibility, we also introduce NIH-TrialBench, a clinician-authored dataset comprising 126 diverse synthetic patient vignettes and matching scenarios from 11 NIH Institutes and Centers. Together, these results support the value of AI to assist clinical trial matching by improving clinician efficiency and identifying frequently overlooked trial opportunities, ultimately helping to expand and accelerate accrual to cancer trials.

Summary

Main Finding

TrialGPT 2.0 is an LLM-based, configurable, and explainable clinical-trial recommendation system that (a) ranks and recommends trials (not just assesses eligibility) using site-specific matching policies and (b) was evaluated across retrospective multicenter and prospective real-world oncology workflows. In retrospective multicenter cohorts (288 de-identified cases, 1,340 patient-trial pairs) it retrieved at least one clinician-recommended trial in its top 10 for ~91% of cases and reduced clinician screening time by 55.0%. In a 6-month prospective Precision Oncology Tumor Board (POTB) deployment (27 live referrals, 339 patient-trial pairs), the system contributed additional trial opportunities missed by routine workflow and expanded patient access to clinical trials by 90.9%. The authors also released NIH-TrialBench, a reproducible benchmark of 126 clinician-authored synthetic vignettes.

Key Points

  • Conceptual advance: shifts from eligibility-only assessment to recommendation-aware matching (accounts for clinical priorities, therapeutic value, workflow constraints).
  • Inputs and configurability:
    • Inputs: patient context (notes or summaries), a predefined local trial corpus, and a local matching policy (e.g., disease/biomarker priorities, preference for therapeutic vs diagnostic trials).
    • Retrieval: pluggable retrieval (lexical + semantic “hybrid-fusion”) for large local corpora; full pass for small corpora.
  • Outputs (inspectable): ranked trials with recommendation categories (Highly Recommended / Possible Match / Low Fit), fit score, confidence, and structured explanations separating eligible reasons, ineligible reasons, missing info, and rationale.
  • Evaluation design:
    • Retrospective multicenter clinician-adjudicated review across five recruitment workflows (OPR, CCF, UIC, UPMC, NCI).
    • Prospective live POTB evaluation embedded in routine tumor-board practice for 6 months.
    • Technical validation on NIH-TrialBench (126 synthetic vignettes) and public benchmarks.
    • Real-world deployment testing (web interface, runtime, reviewer acceptance).
  • Key quantitative results:
    • Retrospective: top-10 recovery of clinician-recommended trials ≈ 91%; clinician screening time reduced by 55.0%.
    • Prospective: TrialGPT 2.0 contributed additional actionable trial options, expanding patient access by 90.9% in the POTB reviews.
    • NIH-TrialBench: 126 vignettes authored by 24 NIH investigators across 11 NIH institutes to enable reproducible evaluation.
  • Implementation: web-based interface with post-ranking filters and local-policy driven sorting; weekly refresh of trial corpus in the POTB setting.

Data & Methods

  • System: TrialGPT 2.0 builds on an earlier TrialGPT LLM framework; adds configurability for local policies and structured, criterion-level assessments. Retrieval uses hybrid lexical/semantic fusion when candidate lists are large (>~1,500).
  • Cohorts:
    • Retrospective: 288 de-identified clinical-note cases spanning five recruitment pathways; local trial lists ranged from 9 to 1,871 trials; total 1,340 reviewed patient-trial pairs.
    • Prospective POTB: 27 live de-identified cases, 339 patient-trial pairs, system outputs reviewed in parallel with human-generated candidate lists; trial corpus refreshed weekly.
    • NIH-TrialBench: 126 synthetic vignettes created by clinicians across 11 NIH Institutes; evaluated against a fixed set of 1,373 candidate trials.
  • Evaluation procedures:
    • Clinician-adjudicated labeling of top-ranked recommendations (top-10 in many cohorts) and exhaustive review for smaller curated lists.
    • Measurements: recovery of clinician-recommended trials in top-K ranks, reduction in time required for clinician screening, whether system-identified trials were adopted into final tumor-board recommendations.
    • Additional technical validation on public eligibility-focused benchmarks.
  • Explainability: Trial-level structured assessments (eligible/ineligible reasons, missing info, rationale) intended for transparent expert review and manual correction.

Implications for AI Economics

  • Direct labor-cost effects:
    • A reported 55% reduction in clinician screening time implies substantial labor cost savings per patient-screening task. For large screening volumes, these savings scale—potentially reducing site staffing costs or enabling staff to screen more candidates per unit time.
  • Trial accrual and failure risk:
    • By expanding identified trial opportunities (prospective increase of ~90.9% in POTB), AI assistance could increase accrual rates, shortening recruitment periods and lowering the probability of early termination due to poor enrollment. Given enrollment shortfalls account for a large share of RCT discontinuations, even moderate accrual improvements can materially reduce cost and time overruns.
  • Cost of drug development and ROI levers:
    • Clinical trials are a major driver of drug-development expense (accounting for over half of total cost in many estimates). Improving recruitment efficiency reduces per-trial duration and site costs and could shrink overall trial budget risk—creating a clear ROI pathway for sponsors, CROs, and institutions adopting matching tools.
  • Market and organizational impacts:
    • Demand for institution-tailored, explainable trial-matching systems is likely to grow—opportunities for vendors, CROs, and hospital informatics teams. Value accrues to large-volume centers and precision-oncology programs where trial portfolios and molecular complexity make manual matching intensive.
  • Deployment considerations with economic consequences:
    • Integration costs: EHR integration, maintenance of local trial lists, periodic model updates, and clinician training represent upfront and ongoing costs.
    • Human-in-the-loop requirement: explainability and clinician adjudication are still required to maintain safety and regulatory compliance; this limits full automation but preserves trust and reduces liability risk.
    • Data & privacy: solutions must navigate clinical data governance; synthetic benchmarks like NIH-TrialBench help reproducible evaluation but do not eliminate integration complexity for real patient data.
  • Risks and negative externalities:
    • False positives/negatives or over-recommendation could lead to wasted screening effort or missed opportunities—economic value depends on high precision in top-ranked recommendations.
    • Unequal adoption could shift accrual toward well-resourced centers, affecting market competition among trial sites and possibly altering site selection economics for sponsors.
  • Recommended economic evaluation metrics and next steps:
    • Track per-site metrics pre/post deployment: screening time per patient, number of patients referred to trials, enrollment rate, time to reach enrollment targets, trial duration and cost per enrolled patient, and trial failure rates due to accrual.
    • Conduct randomized or stepped-wedge trials of adoption at scale to estimate causal effects on accrual, time-to-completion, and cost savings.
    • Model long-run ROI considering deployment and maintenance costs, staff retraining, and expected increases in enrollment and reductions in trial duration/failure rates.
    • Consider pricing models: subscription for institutional deployment, per-screened-patient fees, or vendor integration with CRO workflows.
  • Broader economic policy implications:
    • Regulators and funders could encourage/mandate evaluation standards (transparency, auditability) for matching tools to ensure fair access and reliable benefit measurement.
    • Investment in benchmark datasets (like NIH-TrialBench) and standards will reduce adoption friction and enable comparative cost-effectiveness studies across tools.

Summary: TrialGPT 2.0 demonstrates that configurable, explainable LLM-based trial recommendation systems can materially reduce clinician workload and surface additional trial opportunities in real-world oncology workflows—effects that, if replicated at scale, could lower recruitment-related costs and risks in clinical research. For AI economics, the most important next steps are rigorous, causal evaluations of impact on enrollment and trial cost, and careful accounting of deployment/maintenance costs to establish net ROI.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper reports multicenter retrospective clinician-adjudicated results (288 cases, 1,340 patient-trial pairs), reproducible synthetic vignette benchmarks (126 vignettes), and a prospective 6-month embedded evaluation (27 live tumor-board cases, 339 pairs). These provide real-world and clinician-facing evidence of utility (recall of clinician-recommended trials, screening time reductions, and added trial opportunities). However, there is no randomized or quasi-experimental counterfactual design to rule out selection, reviewer, or temporal biases; the prospective sample is small and oncology-focused, and some evaluation relies on synthetic vignettes. Methods Rigormedium — The study uses multiple complementary evaluation modalities (retrospective multicenter clinician adjudication, prospective embedded workflow assessment, and reproducible vignette benchmarks) and reports relevant performance metrics (top-k recovery, screening time, contribution to final recommendations). However, it lacks randomized allocation or blinded comparison, the prospective evaluation is small (27 cases), clinician review may introduce subjective bias, and external validity beyond participating centers and oncology workflows is uncertain. SampleMulticenter retrospective cohort: 288 de-identified clinical-note cases across five recruitment pathways (NIH Office of Patient Recruitment intake, Cholangiocarcinoma Foundation advocacy intake, UIC tumor-board referrals, UPMC radiation-oncology consults, and NCI referrals) yielding 1,340 reviewed patient-trial pairs; prospective evaluation: 27 live Precision Oncology Tumor Board (POTB) referrals from UIC with 339 patient-trial pairs evaluated over six months; NIH-TrialBench: 126 clinician-authored synthetic vignettes from 11 NIH institutes evaluated against a 1,373-trial search space; additional evaluation on three public benchmarks. Local trial lists per setting ranged from 9 to 1,871 candidate trials. Themeshuman_ai_collab productivity adoption IdentificationNo formal causal identification (no randomization or counterfactual design); evaluation combines retrospective clinician-adjudicated review, technical benchmarking on synthetic and public datasets, and a prospective, non-randomized deployment in a tumor-board workflow where TrialGPT 2.0 outputs were reviewed in parallel with routine human workflows and contribution to final recommendations was recorded. GeneralizabilityPredominantly oncology-focused and concentrated in US academic and government centers — findings may not generalize to non-oncology specialties or community settings., Prospective sample is small (27 cases) and from a single tumor-board; external prospective performance may differ., Performance depends on local trial inventories, search-space scoping, and quality/format of clinical documentation — results may vary across institutions and EHR data quality., Clinician-adjudicated endpoints are subject to subjective judgment and potential reviewer bias; no randomized or blinded comparison to eliminate selection effects., Synthetic vignette benchmarks aid reproducibility but may not capture full complexity of real patient records.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial among its top 10 recommendations for approximately 91% of cases. Output Quality positive Recovery of at least one clinician-recommended clinical trial in the top 10 recommendations
Reading fidelity high
Study strength medium
n=288
approximately 91% of cases
0.48
TrialGPT 2.0 reduced clinician screening time by 55.0% in the retrospective multicenter evaluation. Task Completion Time positive Clinician time spent screening candidate clinical trials
Reading fidelity high
Study strength medium
n=288
55.0% reduction
0.48
During a six-month prospective evaluation embedded in an active precision oncology tumor board, TrialGPT 2.0 contributed additional trial opportunities that were missed by the routine workflow, expanding patient access to clinical trial participation by 90.9%. Consumer Welfare positive Expansion of patient access to potential clinical trial participation through additional trial opportunities
Reading fidelity high
Study strength medium
n=27
90.9% expansion
0.48
The retrospective clinical evaluation covered 1,340 heterogeneous patient-trial pairs from 288 de-identified clinical-note cases across four oncology-specific workflows and one broader NIH referral workflow. Other mixed Coverage and scale of retrospective patient-trial matching evaluation
Reading fidelity high
Study strength high
n=288
1,340 patient-trial pairs from 288 cases
0.8
In the prospective evaluation, TrialGPT 2.0 and clinical reviewers identified candidate trials in parallel, and the combined lists were reviewed by the tumor board to select final trial recommendations. Decision Quality mixed Contribution of AI-identified trials to final clinician-selected recommendations
Reading fidelity high
Study strength medium
n=27
0.48
NIH-TrialBench comprises 126 clinician-authored synthetic patient vignettes and matching scenarios spanning oncology and other disease areas, created by 24 NIH trial investigators from 11 NIH Institutes and Centers. Training Effectiveness positive Availability and breadth of a reproducible clinical-trial-matching evaluation benchmark
Reading fidelity high
Study strength medium
n=126
126 synthetic vignettes; 24 investigators; 11 NIH Institutes and Centers
0.48
TrialGPT 2.0 provides ranked clinical-trial recommendations with structured explanations that separate eligibility evidence, incompatibilities, missing information, and rationale for clinician review. Ai Safety And Ethics positive Inspectability and transparency of AI-generated clinical-trial recommendations
Reading fidelity high
Study strength low
not reported
0.24
TrialGPT 2.0 was deployed in real-world trial-review workflows through a web-based interface. Organizational Efficiency positive Operational deployment of an AI-assisted clinical-trial review system
Reading fidelity high
Study strength low
not reported
0.24

Notes