The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models dramatically lift novices on complex biological tasks: participants with LLM access were over four times more accurate than those limited to web searches and beat experts on multiple benchmarks; nonetheless, users frequently failed to fully extract the models' best answers and safeguards did little to block access to dual-use information.

LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus, Jason Hausenloy, Pedro Medeiros, Nathaniel Li, Aiden Kim, Yury Orlovskiy, Coleman Breen, Bryce Cai, Jasper Götting, Andrew Bo Liu, Samira Nedungadi, Paula Rodriguez, Yannis Yiming He, Mohamed Shaaban, Zifan Wang, Seth Donoughe, Julian Michael · February 26, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Chen Bo Calvin Zhang unresolved corpus identity
  2. Christina Q. Knight unresolved corpus identity
  3. Nicholas Kruus unresolved corpus identity
  4. Jason Hausenloy unresolved corpus identity
  5. Pedro Medeiros unresolved corpus identity
  6. Nathaniel Li unresolved corpus identity
  7. Aiden Kim unresolved corpus identity
  8. Yury Orlovskiy unresolved corpus identity
  9. Coleman Breen unresolved corpus identity
  10. Bryce Cai unresolved corpus identity
  11. Jasper Götting unresolved corpus identity
  12. Andrew Bo Liu unresolved corpus identity
  13. Samira Nedungadi unresolved corpus identity
  14. Paula Rodriguez unresolved corpus identity
  15. Yannis Yiming He unresolved corpus identity
  16. Mohamed Shaaban unresolved corpus identity
  17. Zifan Wang unresolved corpus identity
  18. Seth Donoughe unresolved corpus identity
  19. Julian Michael unresolved corpus identity

Semantic Scholar

Latest observation:

  1. C. Zhang provider ID
  2. Christina Q. Knight provider ID
  3. Nicholas Kruus provider ID
  4. Jason Hausenloy provider ID
  5. Pedro Medeiros provider ID
  6. Nathaniel Li provider ID
  7. Aiden Y Kim provider ID
  8. Yury Orlovskiy provider ID
  9. Coleman Breen provider ID
  10. Bryce Cai provider ID
  11. Jasper Götting provider ID
  12. A. B. Liu provider ID
  13. S. Nedungadi provider ID
  14. Paula Rodriguez provider ID
  15. Yan He provider ID
  16. Mohamed E A Shaaban provider ID
  17. Zifan Wang provider ID
  18. Seth Donoughe provider ID
  19. Julian Michael provider ID
In a randomized multi-benchmark study, novices given LLM access were 4.16 times more accurate on complex biosecurity-relevant tasks than controls using internet resources, often outperforming expert baselines, though standalone LLMs sometimes exceeded assisted users.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurity-relevant task sets. Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). We found that LLM access provided substantial uplift: novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Perhaps surprisingly, standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Overall, LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners, underscoring the need for sustained, interactive uplift evaluations alongside traditional benchmarks.

Summary

Main Finding

LLM access substantially uplifts novices on biosecurity-relevant in silico biology tasks. In a multi-model, multi-benchmark human study, novices given multiple frontier LLMs were on average 4.16× more accurate than internet-only controls (95% CI [2.63, 6.87]). LLM-assisted novices outperformed controls on 7 of 8 benchmarks and exceeded expert baselines on 3 of 4 benchmarks with expert data. However, standalone LLMs often outperformed the assisted humans, suggesting suboptimal human use of models. Most participants (89.6%) reported little difficulty obtaining dual-use–relevant information despite applied safeguards.

Key Points

  • Sample and tasks
    • Participants: STEM cohort N=47 (programming/engineering background) and non‑STEM cohort N=10 (humanities), all classified as novices for complex wet‑lab tasks.
    • Benchmarks: eight biosecurity-relevant tasks (e.g., Virology Capabilities Test, Human Pathogen Capabilities Test, Molecular Biology Capabilities Test, LAB‑Bench, Long‑Form Virology, Agentic Bio‑Capabilities, World Class Biology, Humanity’s Last Exam).
    • Interaction time: allowed extended, iterative work (up to ~13 hours for most involved tasks).
  • Treatment vs Control
    • Treatment: access to multiple frontier LLMs (o3, o4‑mini, Gemini 2.5 Pro, Gemini Deep Research, Claude 3.7 Sonnet, Claude Opus 4 after release, etc.); participants could coordinate across models.
    • Control: internet-only search; LLM features in search disabled.
  • Quantitative outcomes
    • Aggregate uplift: 4.16× higher accuracy for Treatment vs Control (95% CI [2.63, 6.87]).
    • Benchmark coverage: LLM-assisted novices beat controls on 7/8 benchmarks; outperformed expert baselines on 3/4 benchmarks with expert data.
    • Standalone LLMs: frequently scored better than LLM‑assisted novices, implying human factors limit realized model performance.
  • Qualitative/behavioral findings
    • Collected longitudinal data: progress reports, confidence, notes, and interaction logs to study how and when models add value or saturate.
    • Cataloged 28 interaction behaviors (deference, independence, safety-relevant behaviors).
    • 89.6% of participants reported low difficulty in obtaining dual‑use information despite safeguards.
  • Limitations noted by authors
    • Models available changed during the study (Opus 4 released mid‑study).
    • Not double‑blinded; some potential for off‑platform model use (tracked but not airtight).
    • Small-ish overall sample for broad generalization (N=57) though larger than many prior uplift experiments.
    • Some benchmark-specific procedural omissions and question leakages required adaptive mitigation.

Data & Methods

  • Design
    • Non‑STEM cohort: within‑subject alternating Treatment/Control across tasks.
    • STEM cohort: between‑subject assignment to a single condition for coding/agentic tasks.
  • Benchmarks & scoring
    • Mixture of multiple‑choice/multi‑select, short response, and agentic/coding tasks drawn from public and proprietary biosecurity‑relevant benchmarks.
    • LLM baselines: multiple runs per model (e.g., 10 trials per LFV model); refusals scored as zeros in some cases.
    • Expert baselines: practitioners given 15–30 minutes per question, with internet allowed but no LLM assistance.
  • Data collection & safeguards
    • Tracked all on‑platform LLM calls; rate limits for internet‑enabled models (e.g., Gemini Deep Research: 1 request/hour).
    • Participants kept Google Doc notes and periodic “best guess” entries for longitudinal analysis.
  • Analysis
    • Mixed quantitative evaluation (linear and logistic mixed models with participant and question effects, Monte Carlo CIs, Benjamini‑Hochberg FDR control).
    • Qualitative analysis using condition‑blind LLM annotators, embeddings, and regex; cross‑benchmark behavioral coding.
  • Key numeric outcomes highlighted
    • 4.16× uplift (95% CI [2.63, 6.87]).
    • 7/8 benchmarks: Treatment > Control.
    • Treatment > Experts on 3/4 benchmarks with available expert baselines.
    • 89.6% participants reported little difficulty obtaining sensitive information.

Implications for AI Economics

  • Labor and skill diffusion
    • Lowering of task‑specific skill barriers: LLMs can substitute for specialized biological knowledge on many in‑silico tasks, reducing returns to some forms of human expertise (task‑level human capital).
    • Recomposition of labor: value shifts toward roles that verify, interpret, validate, or operationalize model outputs (credentialed oversight, laboratory oversight, auditing), increasing demand and wages for those verification skills while potentially compressing wages for pure task execution.
  • Market structure & product differentiation
    • Multi‑model mosaics matter: access to multiple LLMs materially raises capabilities. Markets may bifurcate into (a) broad-access, higher‑risk multi‑model platforms and (b) safety‑differentiated, gated offerings with stricter controls—supporting premium pricing for safety/assurance.
    • Emergence of adjacent markets: demand for verification services, provenance/audit logs, model‑usage monitoring, and specialized safety layers (technical and legal) is likely to grow.
  • Externalities, risk pricing, and regulation
    • Increased negative externalities: easier access to dual‑use knowledge means private gains from LLM features can generate public harms; traditional market prices will not internalize these risks without regulation or liability.
    • Insurance and financing impacts: heightened dual‑use risk may raise compliance costs, insurance premiums for biotech firms and cloud providers, and influence investor due diligence—especially for startups offering model access to biological actors.
    • Policy levers: access controls (credentialing, zero‑trust APIs), usage quotas, liability rules, and mandatory uplift evaluations could be justified to internalize social costs.
  • Innovation and R&D economics
    • Faster scientific iteration: legitimate R&D can accelerate (lower cost of literature digestion, protocol design), increasing aggregate productivity in biotech but complicating cost–benefit tradeoffs for open access.
    • Potential for misallocated innovation incentives: firms may invest in model access rather than human expertise, altering long‑run investment in training and in‑house R&D capabilities.
  • Modeling and empirical needs
    • Incorporate human–model complementarities into economic models: treat LLMs as capital goods whose marginal productivity depends on user skill; account for learning curves where humans initially underutilize models (observed here) but may improve over time.
    • Endogenous adoption and externalities: models should capture strategic interactions (multi‑model use, verification markets) and dynamic regulatory responses.
  • Short policy recommendations for practitioners and policymakers
    • Fund and require ongoing uplift evaluations that measure human+LLM performance, not just standalone model benchmarks.
    • Consider licensing or credential gating for high‑risk model endpoints and create price/supply mechanisms that reflect social risk (e.g., higher access costs, mandated auditing for biological use‑cases).
    • Support markets for verification, provenance, and incident insurance to internalize risks and create incentives for safer deployment.

Caveats: the study focuses on in silico tasks (non‑wet‑lab execution) and defined “novices” for particular biological tasks; real‑world malicious operationalization still depends on resource, wet‑lab access, and risk of detection. Nonetheless, the findings materially change the economic calculus about who can produce valuable bio‑knowledge and how markets and regulation should respond.

Assessment

Paper Typerct Evidence Strengthmedium — Random assignment across treatment and control supports a causal interpretation of LLM-driven uplift and the study uses multiple benchmarks and models, but key details (sample size, recruitment and randomization checks, blinding, scoring procedures) are not provided here and external/ecological validity is limited to specific biosecurity tasks and model versions. Methods Rigormedium — Strengths include experimental design, multi-model and multi-benchmark coverage, and lengthy task windows; weaknesses include likely unblinded scoring, potential selection bias in novices, unclear sample size and power, limited information on coder reliability and model prompting constraints, and potential deviations from real-world settings. SampleNovice (non-expert) participants were recruited to solve eight complex, biosecurity-relevant task sets with ample time (up to 13 hours for the most involved tasks); treatment groups received access to one or more LLMs (multiple models tested) while controls had internet-only resources; expert baselines were available for four benchmarks for comparison. Themeshuman_ai_collab productivity governance skills_training IdentificationRandomized controlled comparison of novice participants assigned to either LLM access or internet-only access across multiple biosecurity-relevant benchmarks and models, estimating causal uplift via between-group differences (reported effect size 4.16x with 95% CI). GeneralizabilityDomain-specific: tasks focus on biosecurity/biology and may not generalize to other knowledge domains or routine workplace tasks., Novice population: recruited novices may not represent typical workers, students, or malicious actors in the field., Model/version dependency: results depend on specific LLM models, prompts, and configurations tested and will vary as models evolve., Artificial study conditions: long allotted times and experimental setup may differ from real-world, time-constrained decision environments., Limited expert baselines: expert comparisons were available for only half the benchmarks, limiting broader validation., Safety/sandbox differences: safeguards and internet filters used in the study may not match real-world deployments, affecting access to dual-use information.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). Output Quality positive accuracy
Reading fidelity high
Study strength high
4.16 times more accurate (95% CI [2.63, 6.87])
1.0
On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Output Quality positive task performance relative to expert baseline
Reading fidelity high
Study strength medium
not reported
0.6
Standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Output Quality negative comparison of performance (LLM alone vs LLM-assisted human)
Reading fidelity high
Study strength medium
not reported
0.6
Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Ai Safety And Ethics positive self-reported ease of obtaining dual-use-relevant information
Reading fidelity high
Study strength medium
89.6% reported little difficulty
0.6
The study compared novices with LLM access versus internet-only access across eight biosecurity-relevant task sets (multi-model, multi-benchmark human uplift study). Adoption Rate null_result study design / comparison types (LLM vs internet-only) across task sets
Reading fidelity high
Study strength medium
eight biosecurity-relevant task sets
0.6
Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). Task Completion Time null_result time allowed per task
Reading fidelity high
Study strength medium
up to 13 hours
0.6
LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners. Output Quality positive novice ability to perform biological tasks (task performance/accuracy)
Reading fidelity medium
Study strength medium
not reported
0.36
LLMs perform increasingly well on biology benchmarks. Research Productivity positive LLM performance on biology benchmarks
Reading fidelity medium
Study strength speculative
not reported
0.06

Notes