0 cumulative citations
View corpus contextAI is shifting from a niche research tool to an active partner in scientific discovery—foundations models, agentic systems and self-driving labs are speeding research across fields but still face hallucinations, opacity and institutional frictions. The paper maps this diffusion, proposes a typology of research AIs, and warns that technical and social constraints limit their autonomous scientific authority.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper examines the growing role of AI in scientific discovery. It first surveys the rapid rise of AI capabilities, especially in reasoning, abstraction, planning, and long-horizon task execution, before turning to scientometric evidence of AI's diffusion across the sciences. It then proposes a typology of AI systems used in research, ranging from specialized scientific AI through scientific AI assistants and agents to hybrid experimental systems that combine computation and physical experimentation. On this basis, it offers a selective overview of recent achievements in mathematics and computer science, physics, chemistry, the life sciences, and the behavioural and social sciences. It argues that, despite these advances, current systems remain constrained by important technical, epistemic, and institutional limitations, and that their growing use introduces both near-term and longer-term risks. The conclusion further suggests that the advancement of AI in science raises broader questions concerning the division of cognitive labour between human researchers and machines.
Summary
Main Finding
Contemporary AI—driven especially by transformer-based foundation models and multi-agent/hybrid systems—has moved from narrow, auxiliary roles into a significant, heterogeneous set of research participants. AI already performs substantive parts of the scientific pipeline (reasoning, hypothesis generation, experiment design, execution in robotic labs, drafting, coding), but important technical, epistemic, and institutional constraints remain. The development reshapes the division of cognitive labour in science, creating large near-term gains in research productivity and new economic questions and risks about ownership, incentives, measurement, and the distribution of returns from scientific discovery.
Key Points
- Rapid capability growth
- Transformer-based foundation models and LLMs now show improved abstraction, reasoning, planning, and long-horizon task performance; benchmarks (HLE, ARC-AGI) have tracked fast progress.
- Agents can sustain longer multi-step workflows (analogue of “Moore’s Law” for agent task length/complexity proposed).
- Diffusion across disciplines
- AI use in research has expanded rapidly since 2022, becoming pervasive across natural sciences, medicine, social sciences, and humanities.
- Scientometric evidence: large increases in AI-related publications (e.g., Nature portfolio), AI-index reports, and measurable LLM use in manuscript writing (estimates up to ~22% in some fields).
- Typology of AI roles in science (heuristic):
- Specialized Scientific AI — narrowly optimized systems (e.g., AlphaFold).
- Scientific AI Assistants — LLM-based interactive tools for tasks like literature review, coding, writing.
- Scientific AI Agents — multi-agent, more autonomous systems that decompose objectives, run iterative workflows (examples: Sakana’s AI Scientist, Google Co-Scientist, Robin).
- Hybrid AI Experimental Systems — integrated computational–robotic platforms (“self-driving labs”, e.g., ChemAgents, Recursion).
- Achievements and limits
- Demonstrated successes: theorem proving, algorithmic discovery, protein structure prediction, automated chemistry/biotech discovery pipelines.
- Persistent limitations: hallucinations, epistemic opacity, fragility/non-robustness, need for human oversight, benchmark brittleness, remaining gaps to expert-level performance on some frontier tasks.
- Risks and institutional questions
- Epistemic: attribution of discovery, reproducibility, trust in opaque AI outputs.
- Social/institutional: credit allocation, incentives for research, concentration of power (major firms owning foundation models and lab automation), near-term workforce changes.
- Longer-term: governance of increasingly autonomous scientific systems; changes in the meaning of “scientist” vs “operator/supervisor”.
Data & Methods
- Evidence sources
- Literature survey of AI and scientometrics; case studies across disciplines.
- Use of publicly reported benchmark results (e.g., Humanity’s Last Exam (HLE), Abstraction and Reasoning Corpus (ARC-AGI)).
- Scientometric indicators and reports:
- Rise of Generative AI in Science (Ding, Lawson, Shapira 2025).
- AI Index Report 2026 (Stanford).
- Publisher data (counts of “artificial intelligence” occurrences in Nature journals).
- Studies quantifying LLM usage in scientific papers (Liang et al. 2025).
- System-level exemplars and technical papers for concrete demonstrations (AlphaFold, Sakana AI Scientist, ChemAgents, Recursion, Google Co-Scientist, etc.).
- Methods
- Qualitative typology based on roles played in research workflow (not exclusive by architecture).
- Comparative assessment of capabilities via benchmark performance trajectories, task-length metrics, and case outcomes.
- Critical synthesis weighing technical capabilities against epistemic/institutional constraints and risks.
Implications for AI Economics
- AI as a general-purpose technology for R&D
- Broad diffusion suggests AI is acting like a new general-purpose technology (GPT) for science—potentially raising Total Factor Productivity (TFP) in R&D and accelerating innovation.
- Hybrid automation reduces the cost-per-experiment and may compress research cycles, changing returns to scale in discovery-intensive industries (pharma, materials, semiconductors, etc.).
- Division of cognitive labour and labor-market effects
- Task-based substitution and complementarity: routine, codifiable tasks (literature search, coding, drafting, routine analyses, standard assays) are likely to be automated or augmented; creative, supervisory, and high-level integrative tasks may remain complementary to humans.
- Implications for research labor: demand may shift toward roles in AI supervision, experiment design, interpretive judgment, and domain expertise that integrates AI outputs.
- Potential reallocation of returns from labor to capital (owners of models, datasets, lab robotics), with implications for wage inequality among researchers and institutions.
- Measurement and attribution challenges
- Traditional measures of scientific output (publications, patents) may undercount AI contributions; attribution of credit/value to AI-produced or AI-assisted discoveries is ambiguous.
- Productivity metrics need augmentation (e.g., count of experiments per dollar, time-to-discovery, value of enabled follow-on research).
- Market structure and incentives
- Concentration risk: frontier foundation models and lab automation are resource-intensive—likely to be concentrated in large firms or well-funded labs, potentially creating market power over scientific inputs (models, data, robotic platforms).
- Incentive misalignment: proprietary control over pre-trained models or datasets may reduce open science spillovers, slowing wider diffusion of innovation benefits.
- Funding and intellectual property: changes in how discoveries are generated raise questions about patentability, ownership (AI as inventor), and the value capture from AI-generated discoveries.
- Policy and governance implications
- Investment priorities: public funding for open models, shared datasets, and public “self-driving” lab infrastructure can counteract concentration and ensure broader diffusion.
- Regulation and standards: protocols for documentation, reproducibility, provenance, and model explainability will affect scientific validity and market adoption.
- Education and re-skilling: economists and policy makers should support training for researchers in human–AI collaboration, AI oversight, and interdisciplinary skills that complement agents.
- Research agenda for AI economics
- Empirically quantify AI’s causal impact on R&D productivity and time-to-discovery across fields.
- Model the substitution vs complementarity dynamics in scientific labor markets and the resulting distributional effects.
- Analyse market structure: entry barriers to building/operating self-driving labs and foundation models; effects on competition in innovation-intensive sectors.
- Explore incentive mechanisms (public goods provision, IP reform, prizes) to steer AI-enabled discovery toward socially valuable outcomes and safe, verifiable science.
Overall, the paper highlights that AI is already materially altering the economics of scientific discovery—raising opportunities for faster, cheaper innovation but also creating measurement, distributional, governance, and market-structure challenges that merit focused economic research and public policy attention.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AI has already become a significant participant in scientific research and discovery. Research Productivity | positive | AI's integration into scientific research and discovery |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Contemporary frontier AI systems can perform highly autonomous research activities that resemble processes previously carried out exclusively by human scientists, particularly when multiple specialized agents share cognitive labor. Research Productivity | positive | Autonomy in scientific research activities |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AI systems have demonstrated capabilities including theorem proving, algorithmic discovery, and protein structure prediction. Research Productivity | positive | AI performance in scientific discovery tasks |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Frontier models' accuracy on Humanity's Last Exam improved from frequently below 10% in early 2025 evaluations to above 40% for systems such as Gemini 3.1 Pro and GPT-5.5. Research Productivity | positive | Accuracy on Humanity's Last Exam |
Reading fidelity
high
Study strength
medium
|
n=2500
below 10 per cent to above 40 per cent accuracy
|
| AI systems have progressed from completing software-engineering tasks lasting seconds or minutes to undertaking projects requiring several hours of continuous work. Task Completion Time | positive | Duration and complexity of autonomously completed tasks |
Reading fidelity
high
Study strength
medium
|
from seconds or minutes to several hours
|
| AI-related publications in the natural sciences increased by more than one quarter between 2024 and 2025, reaching approximately 80,000 publications in 2025. Adoption Rate | positive | Number of AI-related natural-science publications |
Reading fidelity
high
Study strength
medium
|
more than one quarter increase; approximately 80,000 publications
|
| AI-assisted writing may account for as much as 22% of published work in computer science as of September 2024. Adoption Rate | positive | Share of published scientific work using AI-assisted writing |
Reading fidelity
high
Study strength
medium
|
as much as 22 per cent of published work
|
| The number of papers containing references to artificial intelligence across the Nature journal portfolio increased more than seventy-fold between 2015 and 2025, surpassing 16,000 publications annually. Adoption Rate | positive | Number of AI-related papers in the Nature journal portfolio |
Reading fidelity
high
Study strength
medium
|
more than seventy-fold increase; surpassing 16,000 publications annually
|
| Scientific AI assistants can enhance researcher productivity across activities such as literature review, information retrieval, hypothesis exploration, data analysis, coding, proofreading, critique, and manuscript preparation. Research Productivity | positive | Researcher productivity across scientific-assistance tasks |
Reading fidelity
high
Study strength
low
|
not reported
|
| Current AI systems' scientific successes coexist with hallucinations, epistemic opacity, limited robustness, and a continuing need for human oversight. Ai Safety And Ethics | negative | Reliability and oversight requirements of AI systems used in science |
Reading fidelity
high
Study strength
low
|
not reported
|
| Hybrid AI experimental systems can integrate literature review, experiment design, laboratory operation, data collection, analysis, and refinement of subsequent experimental strategies. Research Productivity | positive | Degree of automation in experimental scientific workflows |
Reading fidelity
high
Study strength
medium
|
not reported
|