0 cumulative citations
View corpus contextAlmost nine in ten open‑access biomedical papers show linguistic traces consistent with LLM assistance by end‑2025, the authors estimate; usage concentrates in Discussion text and is substantially higher in non‑native English countries.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.
Summary
Main Finding
By the end of December 2025, an estimated 89% of open‑access biomedical articles in PubMed Central (PMC) show an excess of LLM‑associated vocabulary consistent with some LLM‑assisted writing or editing. LLM usage rose sharply from early 2023, is concentrated more in Discussion/abstracts than in Methods, and is substantially higher in papers from predominantly non‑native‑English countries.
Key Points
- Core estimate: 0.89 of full papers in the PMC sample show LLM‑associated vocabulary in Dec 2025. (Monthly/yearly estimates rise through 2023–2025.)
- Sectional differences (Dec 2025, entire sections):
- Abstract: 0.68
- Introduction: 0.63
- Methods: 0.54
- Results: 0.58
- Discussion: 0.78
- When using random 255‑word crops (to control for length):
- Discussion: 0.68, Abstract: 0.67, Methods: 0.32 (indicating lower per‑word prevalence in Methods).
- Country differences (2025):
- Highest: South Korea 0.85, China 0.82, Taiwan 0.80.
- Lowest among large contributors: UK 0.28.
- Grouped: native English countries ≈ 0.37 vs non‑native ≈ 0.72.
- The authors validate the estimation method with a simulation (≈100k docs/yr) and recover true β to |estimated − true| < 0.02 across β ∈ [0,1].
- The paper argues earlier frequency‑gap or detector‑based studies tend to undercount LLM usage; survey results (self‑reports) broadly align with higher prevalence.
Data & Methods
- Data
- Source: PMC Open Access Subset (baseline Jan 23, 2026) covering papers through 2025.
- Final sample: 1,194,287 English papers (published 2017–2025) that contain Introduction, Methods, Results, and Discussion sections.
- Top journals in sample include Scientific Reports, Nature Communications, PLoS ONE; ~1.8% preprints.
- Marker words
- Starting list: 379 non‑content words previously identified as having increased usage post‑LLM (e.g., these, potential, delves).
- For robustness, sets of marker words were created by thresholding word frequency (G(T) = words with q < T), and thresholds were optimized per section for 2025.
- Estimation approach
- For each marker word, compute q(t) = fraction of documents containing the word by time t.
- Fit linear regression on 2018–2022 to extrapolate ˆphuman(t), the counterfactual human usage absent LLMs.
- Lower bound logic: for corpus fraction β of LLM‑edited documents, q = (1−β)phuman + β pLLM.
- Using pLLM ≤ 1 yields a usable lower bound β ≥ (q − phuman) / (1 − phuman). Replacing phuman with ˆphuman gives ˆβLB.
- Optimize marker set threshold T to maximize ˆβLB while excluding choices with high standard error (>0.025); take the maximum accepted ˆβLB as the estimate ˆβ.
- Validation and sensitivity
- Simulation with realistic phuman/pLLM ranges and sample sizes replicates β closely (error <0.02).
- Analyses performed at monthly and yearly resolution, per‑section and on fixed‑length random crops (255 words), and by first‑author country affiliation.
- Limitations noted by authors
- Central assumption: linear extrapolation of pre‑LLM trends provides a valid ˆphuman. If humans adopt LLM style (or consciously avoid marker words), estimates can be biased.
- pLLM is not directly observed; the method relies on the conservative pLLM ≤ 1 approximation to produce an estimate interpreted as close to the true value.
- Dataset limited to PMC open‑access subset and to papers with standard sections—may not generalize to all biomedical publishing.
- Marker set and thresholding choices can affect estimates; country‑level estimates used a fixed marker set and may be underestimates.
Implications for AI Economics
- Productivity and labor effects
- Large, rapid adoption (near‑ubiquity by end‑2025 in PMC) implies substantial labor‑augmenting effects in academic writing: time savings, faster submission cycles, and possible increases in output per researcher.
- Potential reallocation of researcher effort from drafting language to other tasks, but also risk of reduced cognitive engagement with writing (creative and critical tasks may see changing skill demands).
- Markets and services
- Demand for traditional language‑editing and translation services may shrink or shift toward hybrid offerings (LLM‑assisted editing, quality control, hallucination mitigation).
- New markets for LLM governance tools (disclosure verification, detector/audit services) and for high‑quality human oversight/editing likely expand.
- Inequality and comparative advantage
- Higher adoption in non‑native‑English countries suggests LLMs reduce language barriers and can raise publication rates from those regions—potentially narrowing one form of global academic inequality.
- Conversely, homogenization of style and reasoning may erode some comparative advantages tied to linguistic or rhetorical diversity.
- Incentives, quality, and externalities
- Faster/cheaper writing may exacerbate publish‑or‑perish incentives, increasing submission volumes and reviewer burden—raising costs for peer review systems and possibly reducing average quality unless review capacity scales.
- Hallucination risks (misattribution, fabricated citations) pose reputational and verification costs that create negative externalities for scientific publishing; this raises expected compliance/enforcement costs for journals and institutions.
- Policy and measurement needs
- High prevalence motivates proactive policy: disclosure norms, editorial guidelines, tooling for provenance and citation checking, and revised evaluation metrics that value reasoning and reproducibility over surface writing.
- Reliable monitoring (like the frequency‑based approach here) is important for policymakers to track adoption, evaluate impacts, and calibrate regulation or funding for mitigation infrastructure.
- Human capital and skill premiums
- Demand may shift toward skills less automatable by LLMs (critical thinking, study design, domain expertise, verification), altering returns to different researcher skills and possibly changing training priorities.
- Welfare tradeoffs
- Net welfare effects are ambiguous: increased inclusivity and efficiency vs. risks to research integrity, diversity, and reviewer/curation capacity. Economic policy should weigh productivity gains against potential increases in monitoring, verification, and remediation costs.
Short summary: the paper documents near‑ubiquitous LLM‑associated language in PMC by late 2025 using a marker‑word, extrapolation‑plus‑thresholding estimator validated by simulation. That prevalence implies large productivity and market effects, distributional shifts across countries and skills, and meaningful policy choices for academic institutions, publishers, and regulators.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| By December 2025, an estimated 89% of full biomedical papers in the analyzed PubMed Central corpus showed signs of LLM-assisted writing or editing. Adoption Rate | positive | Estimated prevalence of LLM-assisted writing or editing in full papers |
Reading fidelity
high
Study strength
medium
|
n=1194287
89%
|
| Estimated LLM usage in full biomedical papers increased from 19% in 2023 to 52% in 2024 and 77% across 2025. Adoption Rate | positive | Estimated annual prevalence of LLM-assisted writing or editing in full papers |
Reading fidelity
high
Study strength
medium
|
n=1194287
19% in 2023, 52% in 2024, and 77% in 2025
|
| In December 2025, estimated LLM usage was higher in 255-word crops from Discussion sections than in crops from Methods sections: 68% versus 32%. Adoption Rate | positive | Estimated LLM-assisted writing or editing in section-level text crops |
Reading fidelity
high
Study strength
medium
|
n=1194287
68% versus 32%
|
| Even when entire Methods sections were analyzed, more than half of them were estimated to show LLM usage by December 2025, with an estimate of 54%. Adoption Rate | positive | Estimated LLM-assisted writing or editing in entire Methods sections |
Reading fidelity
high
Study strength
medium
|
n=1194287
54%
|
| Among the countries with the largest numbers of papers in the dataset, estimated full-paper LLM usage in 2025 was highest in South Korea at 85%, followed by China at 82% and Taiwan at 80%, and lowest in the UK at 28%. Adoption Rate | mixed | Estimated prevalence of LLM-assisted writing or editing by author-affiliation country |
Reading fidelity
high
Study strength
low
|
85% in South Korea, 82% in China, 80% in Taiwan, and 28% in the UK
|
| Estimated LLM usage in 2025 was higher for papers from predominantly non-native-English-speaking countries than for papers from predominantly native-English-speaking countries: 72% versus 37%. Adoption Rate | mixed | Estimated prevalence of LLM-assisted writing or editing by English-language country group |
Reading fidelity
high
Study strength
low
|
72% versus 37%
|
| The proposed estimation procedure recovered the ground-truth LLM-usage rate within 0.02 in all simulated cases with usage rates from 0 to 1. Adoption Rate | null_result | Absolute error of the estimated LLM-usage rate relative to the simulated ground truth |
Reading fidelity
high
Study strength
medium
|
n=100000
|β̂ − β| < 0.02
|
| The paper's LLM-usage estimates are subject to uncertainty because they assume that linear extrapolation from 2018–2022 provides a faithful estimate of human word frequencies in 2025. Ai Safety And Ethics | mixed | Validity and potential bias of estimated LLM-usage prevalence |
Reading fidelity
high
Study strength
medium
|
not reported
|