4 cumulative citations
View corpus contextLarge language models dramatically lift novices on complex biological tasks: participants with LLM access were over four times more accurate than those limited to web searches and beat experts on multiple benchmarks; nonetheless, users frequently failed to fully extract the models' best answers and safeguards did little to block access to dual-use information.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurity-relevant task sets. Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). We found that LLM access provided substantial uplift: novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Perhaps surprisingly, standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Overall, LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners, underscoring the need for sustained, interactive uplift evaluations alongside traditional benchmarks.
Summary
Main Finding
LLM access substantially uplifts novices on biosecurity-relevant in silico biology tasks. In a multi-model, multi-benchmark human study, novices given multiple frontier LLMs were on average 4.16× more accurate than internet-only controls (95% CI [2.63, 6.87]). LLM-assisted novices outperformed controls on 7 of 8 benchmarks and exceeded expert baselines on 3 of 4 benchmarks with expert data. However, standalone LLMs often outperformed the assisted humans, suggesting suboptimal human use of models. Most participants (89.6%) reported little difficulty obtaining dual-use–relevant information despite applied safeguards.
Key Points
- Sample and tasks
- Participants: STEM cohort N=47 (programming/engineering background) and non‑STEM cohort N=10 (humanities), all classified as novices for complex wet‑lab tasks.
- Benchmarks: eight biosecurity-relevant tasks (e.g., Virology Capabilities Test, Human Pathogen Capabilities Test, Molecular Biology Capabilities Test, LAB‑Bench, Long‑Form Virology, Agentic Bio‑Capabilities, World Class Biology, Humanity’s Last Exam).
- Interaction time: allowed extended, iterative work (up to ~13 hours for most involved tasks).
- Treatment vs Control
- Treatment: access to multiple frontier LLMs (o3, o4‑mini, Gemini 2.5 Pro, Gemini Deep Research, Claude 3.7 Sonnet, Claude Opus 4 after release, etc.); participants could coordinate across models.
- Control: internet-only search; LLM features in search disabled.
- Quantitative outcomes
- Aggregate uplift: 4.16× higher accuracy for Treatment vs Control (95% CI [2.63, 6.87]).
- Benchmark coverage: LLM-assisted novices beat controls on 7/8 benchmarks; outperformed expert baselines on 3/4 benchmarks with expert data.
- Standalone LLMs: frequently scored better than LLM‑assisted novices, implying human factors limit realized model performance.
- Qualitative/behavioral findings
- Collected longitudinal data: progress reports, confidence, notes, and interaction logs to study how and when models add value or saturate.
- Cataloged 28 interaction behaviors (deference, independence, safety-relevant behaviors).
- 89.6% of participants reported low difficulty in obtaining dual‑use information despite safeguards.
- Limitations noted by authors
- Models available changed during the study (Opus 4 released mid‑study).
- Not double‑blinded; some potential for off‑platform model use (tracked but not airtight).
- Small-ish overall sample for broad generalization (N=57) though larger than many prior uplift experiments.
- Some benchmark-specific procedural omissions and question leakages required adaptive mitigation.
Data & Methods
- Design
- Non‑STEM cohort: within‑subject alternating Treatment/Control across tasks.
- STEM cohort: between‑subject assignment to a single condition for coding/agentic tasks.
- Benchmarks & scoring
- Mixture of multiple‑choice/multi‑select, short response, and agentic/coding tasks drawn from public and proprietary biosecurity‑relevant benchmarks.
- LLM baselines: multiple runs per model (e.g., 10 trials per LFV model); refusals scored as zeros in some cases.
- Expert baselines: practitioners given 15–30 minutes per question, with internet allowed but no LLM assistance.
- Data collection & safeguards
- Tracked all on‑platform LLM calls; rate limits for internet‑enabled models (e.g., Gemini Deep Research: 1 request/hour).
- Participants kept Google Doc notes and periodic “best guess” entries for longitudinal analysis.
- Analysis
- Mixed quantitative evaluation (linear and logistic mixed models with participant and question effects, Monte Carlo CIs, Benjamini‑Hochberg FDR control).
- Qualitative analysis using condition‑blind LLM annotators, embeddings, and regex; cross‑benchmark behavioral coding.
- Key numeric outcomes highlighted
- 4.16× uplift (95% CI [2.63, 6.87]).
- 7/8 benchmarks: Treatment > Control.
- Treatment > Experts on 3/4 benchmarks with available expert baselines.
- 89.6% participants reported little difficulty obtaining sensitive information.
Implications for AI Economics
- Labor and skill diffusion
- Lowering of task‑specific skill barriers: LLMs can substitute for specialized biological knowledge on many in‑silico tasks, reducing returns to some forms of human expertise (task‑level human capital).
- Recomposition of labor: value shifts toward roles that verify, interpret, validate, or operationalize model outputs (credentialed oversight, laboratory oversight, auditing), increasing demand and wages for those verification skills while potentially compressing wages for pure task execution.
- Market structure & product differentiation
- Multi‑model mosaics matter: access to multiple LLMs materially raises capabilities. Markets may bifurcate into (a) broad-access, higher‑risk multi‑model platforms and (b) safety‑differentiated, gated offerings with stricter controls—supporting premium pricing for safety/assurance.
- Emergence of adjacent markets: demand for verification services, provenance/audit logs, model‑usage monitoring, and specialized safety layers (technical and legal) is likely to grow.
- Externalities, risk pricing, and regulation
- Increased negative externalities: easier access to dual‑use knowledge means private gains from LLM features can generate public harms; traditional market prices will not internalize these risks without regulation or liability.
- Insurance and financing impacts: heightened dual‑use risk may raise compliance costs, insurance premiums for biotech firms and cloud providers, and influence investor due diligence—especially for startups offering model access to biological actors.
- Policy levers: access controls (credentialing, zero‑trust APIs), usage quotas, liability rules, and mandatory uplift evaluations could be justified to internalize social costs.
- Innovation and R&D economics
- Faster scientific iteration: legitimate R&D can accelerate (lower cost of literature digestion, protocol design), increasing aggregate productivity in biotech but complicating cost–benefit tradeoffs for open access.
- Potential for misallocated innovation incentives: firms may invest in model access rather than human expertise, altering long‑run investment in training and in‑house R&D capabilities.
- Modeling and empirical needs
- Incorporate human–model complementarities into economic models: treat LLMs as capital goods whose marginal productivity depends on user skill; account for learning curves where humans initially underutilize models (observed here) but may improve over time.
- Endogenous adoption and externalities: models should capture strategic interactions (multi‑model use, verification markets) and dynamic regulatory responses.
- Short policy recommendations for practitioners and policymakers
- Fund and require ongoing uplift evaluations that measure human+LLM performance, not just standalone model benchmarks.
- Consider licensing or credential gating for high‑risk model endpoints and create price/supply mechanisms that reflect social risk (e.g., higher access costs, mandated auditing for biological use‑cases).
- Support markets for verification, provenance, and incident insurance to internalize risks and create incentives for safer deployment.
Caveats: the study focuses on in silico tasks (non‑wet‑lab execution) and defined “novices” for particular biological tasks; real‑world malicious operationalization still depends on resource, wet‑lab access, and risk of detection. Nonetheless, the findings materially change the economic calculus about who can produce valuable bio‑knowledge and how markets and regulation should respond.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). Output Quality | positive | accuracy |
Reading fidelity
high
Study strength
high
|
4.16 times more accurate (95% CI [2.63, 6.87])
|
| On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Output Quality | positive | task performance relative to expert baseline |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Output Quality | negative | comparison of performance (LLM alone vs LLM-assisted human) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Ai Safety And Ethics | positive | self-reported ease of obtaining dual-use-relevant information |
Reading fidelity
high
Study strength
medium
|
89.6% reported little difficulty
|
| The study compared novices with LLM access versus internet-only access across eight biosecurity-relevant task sets (multi-model, multi-benchmark human uplift study). Adoption Rate | null_result | study design / comparison types (LLM vs internet-only) across task sets |
Reading fidelity
high
Study strength
medium
|
eight biosecurity-relevant task sets
|
| Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). Task Completion Time | null_result | time allowed per task |
Reading fidelity
high
Study strength
medium
|
up to 13 hours
|
| LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners. Output Quality | positive | novice ability to perform biological tasks (task performance/accuracy) |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| LLMs perform increasingly well on biology benchmarks. Research Productivity | positive | LLM performance on biology benchmarks |
Reading fidelity
medium
Study strength
speculative
|
not reported
|