80 cumulative citations
View corpus contextScientists who adopt large language models publish many more preprints — gains of roughly 24–89% depending on field — but much of the extra output is stylistically polished yet substantively weaker. LLM users also draw on a wider, younger literature, forcing journals and funders to rethink how scientific contribution is evaluated.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.
Summary
Main Finding
Large Language Model (LLM) use is already reshaping scientific production: adoption is associated with large increases in paper output (23.7–89.3% depending on field and author background), a breakdown—and in many cases a reversal—of the positive link between linguistic complexity and perceived research quality, and a measurable diversification of the literature researchers access and cite (more books, younger and less-cited documents). The authors argue these shifts will require rethinking evaluation, incentives, and quality-assurance mechanisms in science.
Key Points
-
Scope and headline effects
- Data: 2.1M preprints, 28K referee reports (ICLR‑2024), and 246M document accesses.
- Productivity: LLM adoption associated with increases in submission rates — arXiv +36.2%, bioRxiv +52.9%, SSRN +59.8% (author‑level event‑study estimates; robust across sensitivity checks).
- Heterogeneity: Larger productivity boosts for authors likely to be non‑native English speakers (proxied by names and affiliations). E.g., scholars with Asian names: arXiv +43.0% to bioRxiv +89.3% / SSRN +88.9%; Caucasian names in English‑speaking institutions: +23.7% (arXiv) to +46.2% (SSRN).
-
Writing complexity vs. perceived quality
- LLM‑assisted manuscripts have higher measured linguistic complexity (inverse Flesch reading ease and other lexical/syntactic metrics).
- For human‑written papers, higher complexity correlates with higher publication/peer‑review success (traditional signal).
- For LLM‑assisted papers the relationship reverses: greater complexity associates with worse peer assessments and lower probability of publication. Replicated using 28K ICLR referee reports.
-
Discovery and citation behavior
- Web‑access data (pre/post Bing Chat rollout) show LLM‑aided searchers more often access books (+26.3%), more recent works (median age down ~0.18 years), and less well‑cited documents.
- After adoption, authors cite more diverse sources: +11.9% likelihood of citing books, median cited reference ~0.379 years younger, and no increase in citation impact (mean log citations ~2.34% lower).
- Interpretation: LLMs broaden search/discovery (lower search costs), not simply amplify canonical works.
-
Robustness and replication
- Excluded core AI subfields to avoid confounding from AI research growth.
- Multiple detection thresholds and alternative text features yield consistent patterns.
- Limitations acknowledged: non‑causal observational design, detection based on abstracts, imperfect LLM‑use measurement, adoption endogeneity.
Data & Methods
-
Datasets
- Preprints: arXiv (1.2M), bioRxiv (221K), SSRN (676K); time window Jan 2018–Jun 2024.
- Peer review: ICLR‑2024 referee reports (7,243 submissions, 28K reports).
- Usage logs: 246M arXiv views/downloads with referral sources.
- Citation linkage: OpenAlex and Semantic Scholar (101.6M citation links).
-
LLM‑use detection
- Trained a text‑based detector using token distribution contrasts: human abstracts (pre‑2023) vs. GPT‑3.5turbo‑0125 rewrites to estimate LLM token distribution; compared post‑ChatGPT abstracts to detect probable LLM assistance (thresholding α).
- Sensitivity analyses across thresholds and alternative detectors reported.
-
Empirical strategies
- Author‑level fixed effects event‑study to compare adoption cohorts to similar non‑adopters, excluding core AI areas.
- Heterogeneity analyses by proxying native English status via names and institutional geography.
- Writing complexity measured via additive inverse Flesch Reading Ease; additional lexical/syntactic/morphological features and promotional language used for robustness.
- Publication success proxied by eventual publication in peer‑reviewed venues within observation window; ICLR reviewer scores used as orthogonal quality measure.
- Differences‑in‑differences on web‑access data to assess shifts after Bing Chat (GPT‑4) rollout, using Google referrals as control.
-
Limitations of methods
- Not causal: nonrandom adoption, timing endogeneity, and imperfect measurement of who/how LLMs were used.
- Detection relies on abstracts (not full text); heavy human editing of LLM output can mask use.
- Cannot identify which co‑author used tools or separate writing assistance from other LLM‑enabled tasks (idea generation, coding).
Implications for AI Economics
-
Supply and market structure
- Large productivity gains imply a significant increase in supply of manuscripts. This shifts the “market” for academic attention, increasing congestion and search/selection costs for readers and gatekeepers.
- Geographic redistribution: reduced cost of English communication likely shifts relative market shares toward researchers in non‑native English regions, affecting global competition and comparative advantage in science production.
-
Signaling, incentives, and labor markets
- Traditional signals (linguistic polish, publication counts) become noisy or misleading. This creates adverse selection risks: quantity and surface quality can mask substantive content.
- Hiring, promotion, and funding markets that rely on these signals may misallocate resources unless evaluative criteria adapt.
- Writing labor (editing, translation services) may be partially "automated" or repriced; complementary skills (methodology, experimental rigor, domain expertise) could see increased returns.
-
Information diffusion and knowledge production
- LLMs lower search costs and diversify citations (more books, younger/less‑cited works), potentially accelerating the diffusion of niche or interdisciplinary knowledge and altering citation economies and reputational dynamics.
- Changes in what gets cited will affect long‑run impact metrics and the dynamics of cumulative knowledge (e.g., redistributing attention away from canonical works).
-
Quality assurance and governance
- Increased false signals of quality pressure peer review and editorial systems. Two responses have economic implications:
- Strengthened evaluation (deeper methodological review, reproducibility checks) raises costs per publication (higher reviewer effort, slower throughput).
- Adoption of automated “reviewer agents” or AI evaluators could reallocate reviewer labor but introduces platform/infrastructure externalities and new verification risks.
- Policy instruments (disclosure requirements for AI use, improved detection tools, incentives for replication/robustness) will change cost structures and behavior in research markets.
- Increased false signals of quality pressure peer review and editorial systems. Two responses have economic implications:
-
Measurement, research design, and welfare
- Evaluating social welfare effects requires causal estimates of LLM impacts on research quality, downstream innovation, and welfare (not just output). Key open questions for economic research:
- Do LLM‑enabled increases in output produce more high‑value discoveries per dollar invested?
- How do incentives change for teams vs. individuals, junior vs. senior researchers?
- What are distributional effects across countries, institutions, and demographic groups?
- Evaluating social welfare effects requires causal estimates of LLM impacts on research quality, downstream innovation, and welfare (not just output). Key open questions for economic research:
-
Practical research and policy agenda for economists
- Conduct causal studies (RCTs, instrumenting adoption) to quantify welfare and productivity tradeoffs.
- Model endogenous signaling with noisy surface signals and optimal evaluation policies; analyze market equilibrium under different disclosure/regulation regimes.
- Study labor reallocation and complementarities (which skills gain value as writing is automated).
- Track long‑run citation dynamics and knowledge diffusion under AI‑assisted discovery.
- Design and evaluate incentive and policy interventions: mandatory disclosure, reviewer AI tools, funding/evaluation reforms that weight methodological rigor over surface polish.
In short: LLMs reshape the economics of science by expanding supply, changing comparative advantages, corrupting traditional signals, and altering information flows. Economic analysis and policy should shift from measuring output counts toward measuring and incentivizing substantive quality, rigorous evaluation, and equitable distribution of gains.
Assessment
Claims (5)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production. Research Productivity | positive | paper production (number of papers produced by scientists who adopt LLMs) |
Reading fidelity
high
Study strength
medium
|
n=2100000
23.7-89.3% increase
|
| The magnitude of the increase in paper production among LLM adopters varies substantially by scientific field and author background (range reported 23.7–89.3%). Research Productivity | mixed | heterogeneity in change in paper production across fields and author backgrounds |
Reading fidelity
high
Study strength
medium
|
n=2100000
23.7-89.3% (range across fields and author backgrounds)
|
| LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming. Output Quality | negative | substantive paper quality (as inferred from peer review reports) relative to linguistic complexity |
Reading fidelity
high
Study strength
medium
|
n=28000
|
| LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. Research Productivity | positive | diversity of prior work accessed and cited (types of sources, age of documents, citation counts of cited works) |
Reading fidelity
high
Study strength
medium
|
n=246000000
|
| The observed shifts in scientific production driven by LLM adoption will likely require changes in how journals, funding agencies, and tenure committees evaluate scientific works. Governance And Regulation | mixed | evaluation practices of journals, funders, and tenure committees (institutional evaluation standards) |
Reading fidelity
high
Study strength
speculative
|
not reported
|