An explainable LLM system for oncology trial matching found clinician-recommended trials for the vast majority of retrospective cases and halved screening time, and in a six-month tumor-board pilot it surfaced many trial options missed by routine review, substantially expanding potential trial access.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world deployment. Rather than asking only whether a patient may qualify, the system also assesses which trials warrant further consideration given the patient's current clinical needs and local workflow priorities, and provides structured, inspectable explanations for expert review. Importantly, we evaluated TrialGPT 2.0 retrospectively and prospectively across multiple oncology-focused settings, spanning government, academic cancer-center, patient-advocacy, and NIH referral workflows. In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial in its top 10 recommendations for approximately 91% of cases while reducing clinician screening time by 55.0%. In a six-month prospective evaluation embedded in an active precision oncology tumor board, TrialGPT 2.0 contributed additional trial opportunities missed by the routine workflow, expanding patient access to clinical trial participation by 90.9%. To support scientific reproducibility, we also introduce NIH-TrialBench, a clinician-authored dataset comprising 126 diverse synthetic patient vignettes and matching scenarios from 11 NIH Institutes and Centers. Together, these results support the value of AI to assist clinical trial matching by improving clinician efficiency and identifying frequently overlooked trial opportunities, ultimately helping to expand and accelerate accrual to cancer trials.
Summary
Main Finding
TrialGPT 2.0 is an LLM-based, configurable, and explainable clinical-trial recommendation system that (a) ranks and recommends trials (not just assesses eligibility) using site-specific matching policies and (b) was evaluated across retrospective multicenter and prospective real-world oncology workflows. In retrospective multicenter cohorts (288 de-identified cases, 1,340 patient-trial pairs) it retrieved at least one clinician-recommended trial in its top 10 for ~91% of cases and reduced clinician screening time by 55.0%. In a 6-month prospective Precision Oncology Tumor Board (POTB) deployment (27 live referrals, 339 patient-trial pairs), the system contributed additional trial opportunities missed by routine workflow and expanded patient access to clinical trials by 90.9%. The authors also released NIH-TrialBench, a reproducible benchmark of 126 clinician-authored synthetic vignettes.
Key Points
- Conceptual advance: shifts from eligibility-only assessment to recommendation-aware matching (accounts for clinical priorities, therapeutic value, workflow constraints).
- Inputs and configurability:
- Inputs: patient context (notes or summaries), a predefined local trial corpus, and a local matching policy (e.g., disease/biomarker priorities, preference for therapeutic vs diagnostic trials).
- Retrieval: pluggable retrieval (lexical + semantic “hybrid-fusion”) for large local corpora; full pass for small corpora.
- Outputs (inspectable): ranked trials with recommendation categories (Highly Recommended / Possible Match / Low Fit), fit score, confidence, and structured explanations separating eligible reasons, ineligible reasons, missing info, and rationale.
- Evaluation design:
- Retrospective multicenter clinician-adjudicated review across five recruitment workflows (OPR, CCF, UIC, UPMC, NCI).
- Prospective live POTB evaluation embedded in routine tumor-board practice for 6 months.
- Technical validation on NIH-TrialBench (126 synthetic vignettes) and public benchmarks.
- Real-world deployment testing (web interface, runtime, reviewer acceptance).
- Key quantitative results:
- Retrospective: top-10 recovery of clinician-recommended trials ≈ 91%; clinician screening time reduced by 55.0%.
- Prospective: TrialGPT 2.0 contributed additional actionable trial options, expanding patient access by 90.9% in the POTB reviews.
- NIH-TrialBench: 126 vignettes authored by 24 NIH investigators across 11 NIH institutes to enable reproducible evaluation.
- Implementation: web-based interface with post-ranking filters and local-policy driven sorting; weekly refresh of trial corpus in the POTB setting.
Data & Methods
- System: TrialGPT 2.0 builds on an earlier TrialGPT LLM framework; adds configurability for local policies and structured, criterion-level assessments. Retrieval uses hybrid lexical/semantic fusion when candidate lists are large (>~1,500).
- Cohorts:
- Retrospective: 288 de-identified clinical-note cases spanning five recruitment pathways; local trial lists ranged from 9 to 1,871 trials; total 1,340 reviewed patient-trial pairs.
- Prospective POTB: 27 live de-identified cases, 339 patient-trial pairs, system outputs reviewed in parallel with human-generated candidate lists; trial corpus refreshed weekly.
- NIH-TrialBench: 126 synthetic vignettes created by clinicians across 11 NIH Institutes; evaluated against a fixed set of 1,373 candidate trials.
- Evaluation procedures:
- Clinician-adjudicated labeling of top-ranked recommendations (top-10 in many cohorts) and exhaustive review for smaller curated lists.
- Measurements: recovery of clinician-recommended trials in top-K ranks, reduction in time required for clinician screening, whether system-identified trials were adopted into final tumor-board recommendations.
- Additional technical validation on public eligibility-focused benchmarks.
- Explainability: Trial-level structured assessments (eligible/ineligible reasons, missing info, rationale) intended for transparent expert review and manual correction.
Implications for AI Economics
- Direct labor-cost effects:
- A reported 55% reduction in clinician screening time implies substantial labor cost savings per patient-screening task. For large screening volumes, these savings scale—potentially reducing site staffing costs or enabling staff to screen more candidates per unit time.
- Trial accrual and failure risk:
- By expanding identified trial opportunities (prospective increase of ~90.9% in POTB), AI assistance could increase accrual rates, shortening recruitment periods and lowering the probability of early termination due to poor enrollment. Given enrollment shortfalls account for a large share of RCT discontinuations, even moderate accrual improvements can materially reduce cost and time overruns.
- Cost of drug development and ROI levers:
- Clinical trials are a major driver of drug-development expense (accounting for over half of total cost in many estimates). Improving recruitment efficiency reduces per-trial duration and site costs and could shrink overall trial budget risk—creating a clear ROI pathway for sponsors, CROs, and institutions adopting matching tools.
- Market and organizational impacts:
- Demand for institution-tailored, explainable trial-matching systems is likely to grow—opportunities for vendors, CROs, and hospital informatics teams. Value accrues to large-volume centers and precision-oncology programs where trial portfolios and molecular complexity make manual matching intensive.
- Deployment considerations with economic consequences:
- Integration costs: EHR integration, maintenance of local trial lists, periodic model updates, and clinician training represent upfront and ongoing costs.
- Human-in-the-loop requirement: explainability and clinician adjudication are still required to maintain safety and regulatory compliance; this limits full automation but preserves trust and reduces liability risk.
- Data & privacy: solutions must navigate clinical data governance; synthetic benchmarks like NIH-TrialBench help reproducible evaluation but do not eliminate integration complexity for real patient data.
- Risks and negative externalities:
- False positives/negatives or over-recommendation could lead to wasted screening effort or missed opportunities—economic value depends on high precision in top-ranked recommendations.
- Unequal adoption could shift accrual toward well-resourced centers, affecting market competition among trial sites and possibly altering site selection economics for sponsors.
- Recommended economic evaluation metrics and next steps:
- Track per-site metrics pre/post deployment: screening time per patient, number of patients referred to trials, enrollment rate, time to reach enrollment targets, trial duration and cost per enrolled patient, and trial failure rates due to accrual.
- Conduct randomized or stepped-wedge trials of adoption at scale to estimate causal effects on accrual, time-to-completion, and cost savings.
- Model long-run ROI considering deployment and maintenance costs, staff retraining, and expected increases in enrollment and reductions in trial duration/failure rates.
- Consider pricing models: subscription for institutional deployment, per-screened-patient fees, or vendor integration with CRO workflows.
- Broader economic policy implications:
- Regulators and funders could encourage/mandate evaluation standards (transparency, auditability) for matching tools to ensure fair access and reliable benefit measurement.
- Investment in benchmark datasets (like NIH-TrialBench) and standards will reduce adoption friction and enable comparative cost-effectiveness studies across tools.
Summary: TrialGPT 2.0 demonstrates that configurable, explainable LLM-based trial recommendation systems can materially reduce clinician workload and surface additional trial opportunities in real-world oncology workflows—effects that, if replicated at scale, could lower recruitment-related costs and risks in clinical research. For AI economics, the most important next steps are rigorous, causal evaluations of impact on enrollment and trial cost, and careful accounting of deployment/maintenance costs to establish net ROI.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial among its top 10 recommendations for approximately 91% of cases. Output Quality | positive | Recovery of at least one clinician-recommended clinical trial in the top 10 recommendations |
Reading fidelity
high
Study strength
medium
|
n=288
approximately 91% of cases
|
| TrialGPT 2.0 reduced clinician screening time by 55.0% in the retrospective multicenter evaluation. Task Completion Time | positive | Clinician time spent screening candidate clinical trials |
Reading fidelity
high
Study strength
medium
|
n=288
55.0% reduction
|
| During a six-month prospective evaluation embedded in an active precision oncology tumor board, TrialGPT 2.0 contributed additional trial opportunities that were missed by the routine workflow, expanding patient access to clinical trial participation by 90.9%. Consumer Welfare | positive | Expansion of patient access to potential clinical trial participation through additional trial opportunities |
Reading fidelity
high
Study strength
medium
|
n=27
90.9% expansion
|
| The retrospective clinical evaluation covered 1,340 heterogeneous patient-trial pairs from 288 de-identified clinical-note cases across four oncology-specific workflows and one broader NIH referral workflow. Other | mixed | Coverage and scale of retrospective patient-trial matching evaluation |
Reading fidelity
high
Study strength
high
|
n=288
1,340 patient-trial pairs from 288 cases
|
| In the prospective evaluation, TrialGPT 2.0 and clinical reviewers identified candidate trials in parallel, and the combined lists were reviewed by the tumor board to select final trial recommendations. Decision Quality | mixed | Contribution of AI-identified trials to final clinician-selected recommendations |
Reading fidelity
high
Study strength
medium
|
n=27
|
| NIH-TrialBench comprises 126 clinician-authored synthetic patient vignettes and matching scenarios spanning oncology and other disease areas, created by 24 NIH trial investigators from 11 NIH Institutes and Centers. Training Effectiveness | positive | Availability and breadth of a reproducible clinical-trial-matching evaluation benchmark |
Reading fidelity
high
Study strength
medium
|
n=126
126 synthetic vignettes; 24 investigators; 11 NIH Institutes and Centers
|
| TrialGPT 2.0 provides ranked clinical-trial recommendations with structured explanations that separate eligibility evidence, incompatibilities, missing information, and rationale for clinician review. Ai Safety And Ethics | positive | Inspectability and transparency of AI-generated clinical-trial recommendations |
Reading fidelity
high
Study strength
low
|
not reported
|
| TrialGPT 2.0 was deployed in real-world trial-review workflows through a web-based interface. Organizational Efficiency | positive | Operational deployment of an AI-assisted clinical-trial review system |
Reading fidelity
high
Study strength
low
|
not reported
|