The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models are reshaping the practice of science—speeding literature discovery, ideation, and draft generation—yet their tendency to hallucinate, gaps in evaluation methods, and integrity risks mean they remain powerful assistants rather than reliable substitutes for expert judgment.

Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Steffen Eger, Yong Cao, Jennifer D'Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Moosavi, Wei Zhao, Tristan Miller · September 05, 2026 · ACM Computing Surveys
openalex review_meta n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Steffen Eger provider ID
  2. Yong Cao provider ID
  3. Jennifer D'Souza provider ID
  4. Andreas Geiger provider ID
  5. Christian Greisinger provider ID
  6. Stephanie Gross provider ID
  7. Yufang Hou provider ID
  8. Brigitte Krenn provider ID
  9. Anne Lauscher provider ID
  10. Yizhi Li provider ID
  11. Chenghua Lin provider ID
  12. Nafise Moosavi provider ID
  13. Wei Zhao provider ID
  14. Tristan Miller provider ID

Semantic Scholar

Latest observation:

  1. Steffen Eger provider ID
  2. Yong Cao provider ID
  3. Jennifer D'Souza provider ID
  4. Andreas Geiger provider ID
  5. C. Greisinger provider ID
  6. Stephanie Gross provider ID
  7. Yufang Hou provider ID
  8. Brigitte Krenn provider ID
  9. Anne Lauscher provider ID
  10. Yizhi Li provider ID
  11. Chenghua Lin provider ID
  12. N. Moosavi provider ID
  13. Wei Zhao provider ID
  14. Tristan Miller provider ID
This narrative survey maps how large (multimodal) language models are being used across the scientific lifecycle—search, ideation, experimentation, content and multimodal generation, and peer review—highlighting promising capabilities, evaluation gaps, and ethical risks to research integrity.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to support researchers throughout the scientific lifecycle, including (1) searching for relevant literature, (2) generating research ideas and conducting experiments, (3) producing text-based content, (4) creating multimodal artifacts such as figures and diagrams, and (5) evaluating scientific work, as in peer review. In this survey, we provide a curated overview of literature representative of the core techniques, evaluation practices, and emerging trends in AI-assisted scientific discovery. Across the five tasks outlined above, we discuss datasets, methods, results, evaluation strategies, limitations, and ethical concerns, including risks to research integrity through the misuse of generative models. We aim for this survey to serve both as an accessible, structured orientation for newcomers to the field, as well as a catalyst for new AI-based initiatives and their integration into future “AI4Science” systems.

Summary

Main Finding

Large multimodal language models (LLMs) and an ecosystem of AI tools are already reshaping the scientific research lifecycle — from literature search through idea generation, experimentation, multimodal content creation, and evaluation/peer review. Capabilities vary by task: literature search and summarization tools are relatively mature and widely usable; text generation is highly usable but suffers from factuality issues; ideation and automated experimentation show high novelty but low reliability; multimodal generation is promising visually but limited in accuracy; automated peer review is useful as an assistive tool but unreliable as a standalone evaluator. Major unresolved problems are hallucination, bias, weak evaluation standards, threats to research integrity (plagiarism, fake science), and environmental/externality costs.

Key Points

  • Scope: The survey adopts a workflow-centric view of AI-assisted science, organizing literature around five core tasks: (1) literature search & summarization, (2) idea & hypothesis generation + experimentation, (3) text-based content generation, (4) multimodal content generation (figures, tables, slides), and (5) evaluation & peer review.
  • Representative tools and approaches: Elicit, ORKG ASK, SciSpace, NotebookLM, ChatGPT/Gemini/Claude, DeepSeek, Connected Papers, CiteSpace, AutomaTikZ/DeTikZify, The AI Scientist, and many RAG (retrieval-augmented generation) systems.
  • Core technical patterns:
    • Retrieval-augmented generation (RAG) to ground LLM outputs in documents/knowledge graphs.
    • Knowledge graphs and graph-based bibliometrics for structured queries and trend analysis.
    • Multimodal models for figures/tables and pipelines for code/experiment generation.
    • Human-in-the-loop workflows to mitigate model failures.
  • Evaluation gaps:
    • Lack of standardized, reliable benchmarks across many tasks (especially ideation, experiment planning, figure generation, peer review).
    • Heavy reliance on ad hoc human evaluation; automated metrics inadequate for factuality/truth.
  • Ethical and integrity concerns:
    • Hallucination and factual inconsistency that can propagate false claims.
    • Risks of fake or fraudulent science, plagiarism, questionable authorship.
    • Bias amplification and unequal access (data and compute barriers).
    • Environmental cost of large models and infrastructure.
  • Trends and ecosystem dynamics:
    • Rapid cross-discipline uptake (citation growth of LLMs across 22 non-CS fields).
    • Emergence of specialized AI4Science venues and tooling.
    • Growing interest in integrated AI-assisted research workflows but with fragmentary evaluation and governance.

Data & Methods

  • Data sources surveyed:
    • Open repositories (arXiv, PubMed Central), subscription archives (ScienceDirect), aggregator indices (Semantic Scholar, CORE), preprint servers (bioRxiv), datasets and code repositories (Zenodo, Dryad, GitHub), and domain-specific collections.
    • Scholarly metadata and citation networks employed for graph-based tools (Connected Papers, CiteSpace).
  • Methods summarized:
    • Retrieval & semantic search: dense embeddings, vector search, and RAG pipelines to retrieve and ground content.
    • Knowledge graphs: structured extraction of contributions (claims, method/results) to enable structured QA and comparisons (ORKG ASK).
    • Paper chat/Q&A: chunking PDFs, indexing, then LLM prompting over retrieved chunks (NotebookLM, ChatPDF).
    • Automated experimentation: pipelines that generate code, workflows or propose experiments; some use symbolic/simulated environments (The AI Scientist-style approaches).
    • Text generation: fine-tuning and prompting strategies for titles, abstracts, related work, and rewriting; control for style and formatting remains active research.
    • Multimodal generation: model-conditioned figure/table generation, conversion from LaTeX figures (e.g., DeTikZify, AutomaTikZ).
    • Evaluation methods: human judgment, task-specific benchmarks where available, qualitative case studies; frequent use of small-scale user studies; large-scale, rigorous benchmarks remain scarce.
  • Survey methodology: narrative (not systematic) selection of representative and influential works via expert curation, citation chaining, and targeted keyword searches — emphasis on conceptual coherence and cross-task comparability rather than exhaustiveness.

Implications for AI Economics

The survey points to multiple concrete implications and research opportunities for economists studying AI and innovation:

  • Productivity and output measurement

    • Potential productivity gains in research (faster literature review, draft generation, ideation) warrant rigorous measurement. Economists should quantify effects on publication rates, time-to-result, and quality-adjusted outputs.
    • Recommended empirical approaches: difference-in-differences on tool rollouts, field experiments with research groups, panel studies of grant/productivity pre/post adoption, and task-level time-use studies.
  • Labor market effects & complementarities

    • AI is likely to augment some research tasks while displacing routine tasks (e.g., summarization, drafting). Study heterogenous impacts across roles (principal investigators, postdocs, technicians, research software engineers).
    • Examine reallocation of time to higher-value tasks (conceptual work, experimental design) vs. deskilling risks.
  • Distributional and access effects

    • Tools may lower barriers for non-native English speakers and under-resourced institutions (increasing inclusivity), but compute/data barriers could concentrate advantages in well-funded labs.
    • Study adoption heterogeneity across countries, disciplines, and institution types and resulting inequality in research outcomes.
  • Market structure, firms, and platform economics

    • Emergence of specialized AI4Science tools suggests platform dynamics: bundling, data lock-in, two-sided markets (researchers vs. publishers/data providers).
    • Investigate pricing models (subscription vs. freemium vs. API), strategic entry by large LLM providers, and the role of proprietary vs. open-data ecosystems.
  • Information quality, externalities, and trust

    • Hallucination and misuse generate negative externalities (misinformation, fake science). Economists can design models of incentives for verification services, credentialing, and platform moderation.
    • Value and market for verification/audit services (automated claim-checkers, reproducibility-as-a-service) is likely to grow — evaluate willingness-to-pay and regulatory roles.
  • Intellectual property, attribution, and incentives

    • Authorship, citation, and IP norms may shift with AI-assisted contributions. Study how credit allocation and incentive structures for co-authorship or AI-assisted outputs evolve and affect collaboration and innovation.
  • Public policy and regulation

    • Need cost–benefit analyses of regulation that balances innovation acceleration with integrity risks and environmental costs.
    • Study effects of policy interventions (disclosure requirements, certification of AI-assisted outputs, funding for open datasets) on research productivity and welfare.
  • Environmental and compute-cost accounting

    • Incorporate model-training and inference carbon costs into social accounting of scientific production. Analyze trade-offs between compute-intensive automation and gains in research efficiency.
  • Data & empirical strategies for economists

    • New data sources: usage logs from AI platforms (Elicit, SciSpace), API call traces, download/citation patterns, GitHub code and workflow repositories, publisher metadata, and preprint timelines.
    • Methods: causal inference (rollouts, IVs), structural models of task allocation and investment in compute/data, multi-armed field trials with institutions, and cost-benefit models including environmental externalities.

Suggested near-term research questions - How much does AI-assisted tooling change individual researcher productivity and the distribution of outputs across career stages? - Do AI tools increase the speed of cumulative science (e.g., faster follow-up studies, replication) or primarily produce more low-quality outputs? - What market structures emerge for AI4Science platforms, and how do data-access constraints shape competition? - How effective are disclosure, certification, or auditing policies at preserving research integrity without stifling productivity?

Overall, this survey shows AI models are transforming scientific production in economically meaningful ways — creating new market opportunities, altering task allocation and incentives, and generating externalities that merit careful empirical and policy-focused economic analysis.

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a narrative survey synthesizing existing literature and tools rather than presenting new empirical causal evidence or identification strategies. Methods Rigormedium — The authors use a transparent, narrative survey approach: assembling candidate works from co-authors' seed papers, forward/backward citation chaining, and keyword searches, and selecting representative studies based on relevance, venue reputation, and impact indicators. This is appropriate for a broad, rapidly evolving field but is not a systematic review and therefore is subject to selection bias and incomplete coverage. SampleNarrative synthesis of published literature, tools, datasets, and benchmarks related to LLMs and multimodal models for scientific workflows; selection was via seed papers from co-authors, forward/backward citation analysis, and targeted keyword searches across major scholarly databases, with coverage up to early 2026; no original experiments or primary data collection. Themeshuman_ai_collab productivity adoption innovation governance GeneralizabilityNot exhaustive — narrative selection may omit relevant papers, especially outside co-authors' domains., Likely English-language and CS/NLP–heavy bias; coverage of domain-specific scientific fields may be uneven., Time-limited snapshot (up to ~2025/early-2026) and subject to rapid change in LLM capabilities and tools., Findings are descriptive and conceptual rather than causal, limiting direct transfer to quantitative policy or economic inference.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Citations of large language models increased rapidly between 2018 and 2024 across 22 non-computer-science disciplines. Adoption Rate positive Growth in scientific use or attention to large language models, measured through citations
Reading fidelity high
Study strength medium
n=148000
0.24
Researchers widely expect AI use to become mainstream in scientific practice within the next two years, although current use is often limited to writing assistance. Adoption Rate mixed Researchers' expectations and current use of AI in scientific practice
Reading fidelity high
Study strength medium
not reported
0.24
AI systems are beginning to support multiple stages of the scientific research lifecycle, including literature search, experimentation, multimodal content generation, and evaluation of scientific outputs. Organizational Efficiency positive Breadth of AI support across scientific research tasks
Reading fidelity high
Study strength medium
not reported
0.24
AI-assisted research tools have the potential to accelerate scientific discovery and streamline the documentation and communication of results. Research Productivity positive Speed and efficiency of scientific discovery, documentation, and communication
Reading fidelity high
Study strength speculative
not reported
0.04
AI-enhanced literature-search tools can help researchers compare and interpret scientific literature more efficiently by summarizing and categorizing study outcomes, methodologies, and limitations. Research Productivity positive Efficiency of literature comparison and interpretation
Reading fidelity high
Study strength medium
not reported
0.24
Knowledge-graph-based scientific question-answering systems can provide more interpretable and verifiable answers than traditional LLM-based question-answering systems. Decision Quality positive Interpretability and verifiability of answers to scientific questions
Reading fidelity high
Study strength low
not reported
0.12
Current AI tools for scientific research exhibit hallucination, bias, limited reasoning abilities, and substantial environmental costs, while evaluation mechanisms remain underdeveloped. Ai Safety And Ethics negative Reliability, fairness, reasoning capability, environmental impact, and evaluability of AI research tools
Reading fidelity high
Study strength medium
not reported
0.24
AI-generated scientific content creates risks of fake science, plagiarism, and erosion of research integrity when human oversight is diminished. Ai Safety And Ethics negative Scientific integrity and trustworthiness of research outputs
Reading fidelity high
Study strength medium
not reported
0.24
AI-assisted scientific experimentation is at an early stage and lacks reliability, with systems often making critical errors and struggling to identify flaws. Error Rate negative Reliability and error rate of AI-assisted scientific experimentation
Reading fidelity high
Study strength medium
not reported
0.24
Automated peer-review systems show promise as assistants but are unreliable as standalone evaluators. Decision Quality mixed Reliability and usefulness of AI-assisted scientific evaluation and peer review
Reading fidelity high
Study strength medium
not reported
0.24

Notes