The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Scientists who adopt large language models publish many more preprints — gains of roughly 24–89% depending on field — but much of the extra output is stylistically polished yet substantively weaker. LLM users also draw on a wider, younger literature, forcing journals and funders to rethink how scientific contribution is evaluated.

Scientific production in the era of Large Language Models
Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin · January 19, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Keigo Kusumegi unresolved corpus identity
  2. Xinyu Yang unresolved corpus identity
  3. Paul Ginsparg unresolved corpus identity
  4. Mathijs de Vaan unresolved corpus identity
  5. Toby Stuart unresolved corpus identity
  6. Yian Yin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Keigo Kusumegi provider ID
  2. Xinyu Yang provider ID
  3. P. Ginsparg provider ID
  4. M. D. Vaan provider ID
  5. Toby Stuart provider ID
  6. Yian Yin provider ID
Researchers who adopt LLMs produce substantially more preprints (23.7–89.3% increases) but produce manuscripts that are linguistically more complex yet substantively weaker, while citing a broader and younger set of prior work.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.

Summary

Main Finding

Large Language Model (LLM) use is already reshaping scientific production: adoption is associated with large increases in paper output (23.7–89.3% depending on field and author background), a breakdown—and in many cases a reversal—of the positive link between linguistic complexity and perceived research quality, and a measurable diversification of the literature researchers access and cite (more books, younger and less-cited documents). The authors argue these shifts will require rethinking evaluation, incentives, and quality-assurance mechanisms in science.

Key Points

  • Scope and headline effects

    • Data: 2.1M preprints, 28K referee reports (ICLR‑2024), and 246M document accesses.
    • Productivity: LLM adoption associated with increases in submission rates — arXiv +36.2%, bioRxiv +52.9%, SSRN +59.8% (author‑level event‑study estimates; robust across sensitivity checks).
    • Heterogeneity: Larger productivity boosts for authors likely to be non‑native English speakers (proxied by names and affiliations). E.g., scholars with Asian names: arXiv +43.0% to bioRxiv +89.3% / SSRN +88.9%; Caucasian names in English‑speaking institutions: +23.7% (arXiv) to +46.2% (SSRN).
  • Writing complexity vs. perceived quality

    • LLM‑assisted manuscripts have higher measured linguistic complexity (inverse Flesch reading ease and other lexical/syntactic metrics).
    • For human‑written papers, higher complexity correlates with higher publication/peer‑review success (traditional signal).
    • For LLM‑assisted papers the relationship reverses: greater complexity associates with worse peer assessments and lower probability of publication. Replicated using 28K ICLR referee reports.
  • Discovery and citation behavior

    • Web‑access data (pre/post Bing Chat rollout) show LLM‑aided searchers more often access books (+26.3%), more recent works (median age down ~0.18 years), and less well‑cited documents.
    • After adoption, authors cite more diverse sources: +11.9% likelihood of citing books, median cited reference ~0.379 years younger, and no increase in citation impact (mean log citations ~2.34% lower).
    • Interpretation: LLMs broaden search/discovery (lower search costs), not simply amplify canonical works.
  • Robustness and replication

    • Excluded core AI subfields to avoid confounding from AI research growth.
    • Multiple detection thresholds and alternative text features yield consistent patterns.
    • Limitations acknowledged: non‑causal observational design, detection based on abstracts, imperfect LLM‑use measurement, adoption endogeneity.

Data & Methods

  • Datasets

    • Preprints: arXiv (1.2M), bioRxiv (221K), SSRN (676K); time window Jan 2018–Jun 2024.
    • Peer review: ICLR‑2024 referee reports (7,243 submissions, 28K reports).
    • Usage logs: 246M arXiv views/downloads with referral sources.
    • Citation linkage: OpenAlex and Semantic Scholar (101.6M citation links).
  • LLM‑use detection

    • Trained a text‑based detector using token distribution contrasts: human abstracts (pre‑2023) vs. GPT‑3.5turbo‑0125 rewrites to estimate LLM token distribution; compared post‑ChatGPT abstracts to detect probable LLM assistance (thresholding α).
    • Sensitivity analyses across thresholds and alternative detectors reported.
  • Empirical strategies

    • Author‑level fixed effects event‑study to compare adoption cohorts to similar non‑adopters, excluding core AI areas.
    • Heterogeneity analyses by proxying native English status via names and institutional geography.
    • Writing complexity measured via additive inverse Flesch Reading Ease; additional lexical/syntactic/morphological features and promotional language used for robustness.
    • Publication success proxied by eventual publication in peer‑reviewed venues within observation window; ICLR reviewer scores used as orthogonal quality measure.
    • Differences‑in‑differences on web‑access data to assess shifts after Bing Chat (GPT‑4) rollout, using Google referrals as control.
  • Limitations of methods

    • Not causal: nonrandom adoption, timing endogeneity, and imperfect measurement of who/how LLMs were used.
    • Detection relies on abstracts (not full text); heavy human editing of LLM output can mask use.
    • Cannot identify which co‑author used tools or separate writing assistance from other LLM‑enabled tasks (idea generation, coding).

Implications for AI Economics

  • Supply and market structure

    • Large productivity gains imply a significant increase in supply of manuscripts. This shifts the “market” for academic attention, increasing congestion and search/selection costs for readers and gatekeepers.
    • Geographic redistribution: reduced cost of English communication likely shifts relative market shares toward researchers in non‑native English regions, affecting global competition and comparative advantage in science production.
  • Signaling, incentives, and labor markets

    • Traditional signals (linguistic polish, publication counts) become noisy or misleading. This creates adverse selection risks: quantity and surface quality can mask substantive content.
    • Hiring, promotion, and funding markets that rely on these signals may misallocate resources unless evaluative criteria adapt.
    • Writing labor (editing, translation services) may be partially "automated" or repriced; complementary skills (methodology, experimental rigor, domain expertise) could see increased returns.
  • Information diffusion and knowledge production

    • LLMs lower search costs and diversify citations (more books, younger/less‑cited works), potentially accelerating the diffusion of niche or interdisciplinary knowledge and altering citation economies and reputational dynamics.
    • Changes in what gets cited will affect long‑run impact metrics and the dynamics of cumulative knowledge (e.g., redistributing attention away from canonical works).
  • Quality assurance and governance

    • Increased false signals of quality pressure peer review and editorial systems. Two responses have economic implications:
      • Strengthened evaluation (deeper methodological review, reproducibility checks) raises costs per publication (higher reviewer effort, slower throughput).
      • Adoption of automated “reviewer agents” or AI evaluators could reallocate reviewer labor but introduces platform/infrastructure externalities and new verification risks.
    • Policy instruments (disclosure requirements for AI use, improved detection tools, incentives for replication/robustness) will change cost structures and behavior in research markets.
  • Measurement, research design, and welfare

    • Evaluating social welfare effects requires causal estimates of LLM impacts on research quality, downstream innovation, and welfare (not just output). Key open questions for economic research:
      • Do LLM‑enabled increases in output produce more high‑value discoveries per dollar invested?
      • How do incentives change for teams vs. individuals, junior vs. senior researchers?
      • What are distributional effects across countries, institutions, and demographic groups?
  • Practical research and policy agenda for economists

    • Conduct causal studies (RCTs, instrumenting adoption) to quantify welfare and productivity tradeoffs.
    • Model endogenous signaling with noisy surface signals and optimal evaluation policies; analyze market equilibrium under different disclosure/regulation regimes.
    • Study labor reallocation and complementarities (which skills gain value as writing is automated).
    • Track long‑run citation dynamics and knowledge diffusion under AI‑assisted discovery.
    • Design and evaluate incentive and policy interventions: mandatory disclosure, reviewer AI tools, funding/evaluation reforms that weight methodological rigor over surface polish.

In short: LLMs reshape the economics of science by expanding supply, changing comparative advantages, corrupting traditional signals, and altering information flows. Economic analysis and policy should shift from measuring output counts toward measuring and incentivizing substantive quality, rigorous evaluation, and equitable distribution of gains.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Very large, multi-source datasets (2.1M preprints, 28K peer reviews, 246M access events) and longitudinal comparisons strengthen causal claims, but reliance on non-random, self-reported or inferred LLM adoption, possible selection into adoption, measurement issues for 'quality' and 'LLM use', and potential unobserved confounders limit confidence in a clean causal interpretation. Methods Rigormedium — The study uses rich, large-scale observational data and plausible quasi-experimental designs (event-study/DID, fixed effects, matching, validation with peer reviews), but critical threats remain (selection bias, measurement error in LLM usage and quality proxies, heterogeneous treatment timing, and limited external validation of substantive quality), and no randomized assignment is available. SampleAggregate of 2.1 million preprints across multiple scientific fields (likely from large preprint servers), 28,000 peer-review reports used to validate manuscript quality assessments, and 246 million recorded online accesses to scientific documents; authors are stratified by field and background and LLM use is measured via self-reporting, metadata, or automated detection of LLM-generated text (paper does not report randomized assignment). Time window covers the recent rapid adoption period of LLMs (exact dates not specified). Themesproductivity human_ai_collab adoption innovation IdentificationAuthor-level longitudinal comparison: the paper appears to identify causal effects by comparing researchers before and after reported LLM adoption to contemporaneous non-adopters (an event-study / difference-in-differences framework), with controls for author and field fixed effects, time trends, and observable covariates; robustness checks reportedly include matching and falsification/placebo tests and validation of quality measures with peer-review reports and access logs. Key limitation: adoption is not randomized and likely correlated with unobserved productivity shocks and measurement error in LLM-use indicators. GeneralizabilityPreprints and access logs may not represent peer-reviewed, final published literature, LLM use measurement likely relies on self-report or text inference, risking misclassification, Early adopters differ from later adopters; effects may reflect selection by ambitious or better-resourced scientists, Findings may be biased toward English-language outputs and fields with strong preprint cultures (e.g., physics, CS, bio), Rapid evolution of LLM capability means effects may change over time, Peer-review subsample (28K reports) is much smaller and may not be representative of all journals or disciplines

Claims (5)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production. Research Productivity positive paper production (number of papers produced by scientists who adopt LLMs)
Reading fidelity high
Study strength medium
n=2100000
23.7-89.3% increase
0.48
The magnitude of the increase in paper production among LLM adopters varies substantially by scientific field and author background (range reported 23.7–89.3%). Research Productivity mixed heterogeneity in change in paper production across fields and author backgrounds
Reading fidelity high
Study strength medium
n=2100000
23.7-89.3% (range across fields and author backgrounds)
0.48
LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming. Output Quality negative substantive paper quality (as inferred from peer review reports) relative to linguistic complexity
Reading fidelity high
Study strength medium
n=28000
0.48
LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. Research Productivity positive diversity of prior work accessed and cited (types of sources, age of documents, citation counts of cited works)
Reading fidelity high
Study strength medium
n=246000000
0.48
The observed shifts in scientific production driven by LLM adoption will likely require changes in how journals, funding agencies, and tenure committees evaluate scientific works. Governance And Regulation mixed evaluation practices of journals, funders, and tenure committees (institutional evaluation standards)
Reading fidelity high
Study strength speculative
not reported
0.08

Notes