6 cumulative citations
View corpus contextThe spread of large language models is changing how grant ideas are framed and selected: proposers who use LLMs produce less semantically distinctive submissions and—at NIH but not NSF—are more likely to win funding and generate early, modest-impact papers.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly diffusing into scientific practice, holding substantial promise while raising widespread concerns. Despite growing attention to AI use in scientific writing and evaluation, little is known about how the rise of LLMs is reshaping the public funding landscape. Here, we examine LLM involvement at key stages of the federal funding pipeline by combining two complementary data sources: confidential National Science Foundation (NSF) and National Institutes of Health (NIH) proposal submissions from two large US R1 universities, including funded, unfunded, and pending proposals, and the full population of publicly released NSF and NIH awards. We find that LLM use rises sharply beginning in 2023 and exhibits a bimodal distribution, indicating a clear split between minimal and substantive use. Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency. The consequences of this shift are agency-dependent. LLM use is positively associated with proposal success and higher early-stage publication output at NIH, whereas no comparable associations are observed at NSF. Notably, the productivity gains at NIH are concentrated in non-hit papers rather than the most highly cited work. Together, these findings provide large-scale evidence that the rise of LLMs is reshaping how scientific ideas are positioned, selected, and translated into publicly funded research, with implications for portfolio governance, research diversity, and the long-run impact of science.
Summary
Main Finding
The rapid adoption of large language models (LLMs) in federal grant writing since 2023 is detectable and uneven (bimodal). Higher LLM involvement is consistently associated with proposals and awards that are semantically closer to recently funded work (lower distinctiveness). Consequences differ by agency: at NIH higher LLM involvement is associated with greater proposal success (~+4 percentage points) and ~+5% more early-stage publications (driven by non-hit papers), whereas at NSF no systematic association with funding success or publication output is observed.
Key Points
- Temporal pattern
- Detectable LLM use rose sharply beginning in 2023 (coinciding with ChatGPT availability).
- Individual-grant LLM involvement (α) shows a bimodal distribution in 2023–2025: one mode near zero, another around ~10–15% α.
- Positioning of ideas
- Greater LLM involvement is associated with lower semantic distinctiveness (proposals/awards move closer to the agency’s recent funded portfolio).
- Magnitude: moving an award from the 25th to the 75th percentile of α corresponds to ~5 percentage-point decrease in distinctiveness for NSF and ~4 points for NIH (within-year percentile rank).
- Rewriting 2021 abstracts with an LLM did not materially change distinctiveness measures, suggesting substantive convergence—not only surface-level stylistic smoothing.
- Selection and outputs (agency-dependent)
- NSF: no statistically significant relationship between α and proposal success or publication counts.
- NIH: higher α is associated with higher funding probability (~+4 percentage points from 25th→75th α) and ~+5% more early publications for grants (2023–2024 cohorts).
- The NIH publication gain is concentrated in non-hit papers; associations with top 5% (or top 1%) high-impact papers are not significant.
- Robustness
- Results robust to controls (grant-year, field, investigator fixed effects; funding amount).
- LLM-associated proposals score as more readable (Flesch reading ease), and removing promotional words does not change the LLM-use signal.
Data & Methods
- Data
- Private submissions: complete NSF and NIH proposal abstracts (funded, unfunded, pending) from two large U.S. R1 universities (request start dates 2021–2025).
- Public awards: full population of publicly released NSF and NIH award abstracts (same period) plus linked publications for awards.
- Measuring LLM involvement (α)
- Based on Liang et al. detection approach: estimate word distributions for human-written text (public 2021 abstracts) and for LLM-generated text (GPT-3.5-turbo-0125 rewriting of the same abstracts).
- Applied detection at corpus level (3-month rolling windows) and per-grant level to get α (fraction of LLM-modified sentences).
- Semantic distinctiveness
- Embed abstracts using SPECTER2 embeddings and compute average cosine distance to all agency-funded abstracts in the prior year.
- Convert distances to within-year percentile ranks (higher = more distinctive).
- Statistical analysis
- Regressions of distinctiveness, funding success, and publication outcomes on grant-level α.
- All regressions include start-year, field, and investigator (PI/co-PI) fixed effects and controls for requested/awarded funding amount.
- Analyses exploit within-investigator variation in α (estimates are conditional correlations, not causal effects).
- Robustness checks
- Rewrote 2021 abstracts with LLMs to test whether surface-level changes explain distinctiveness shifts.
- Evaluated reading-ease (Flesch) and removed promotional language; results held.
Implications for AI Economics
- Upstream effects of LLMs matter: LLMs are already reshaping the funding pipeline (positioning and selection), not just downstream manuscript production.
- Tradeoff: individual-level gains vs. system-level diversity
- LLMs can raise submission quality/readability and may increase funded outputs (NIH), but they also appear to pull proposals toward established funded norms, reducing semantic novelty. This suggests a shift from exploration toward exploitation in public research portfolios.
- Agency heterogeneity matters for outcomes and policy
- Differences between NIH and NSF outcomes indicate that institutional review processes, norms, and evaluation criteria interact with LLM effects. Policy responses likely need to be agency-specific.
- Implications for portfolio governance and evaluation
- Funders should monitor LLM use and its impact on portfolio diversity and long-run innovation. Consider requiring disclosure of generative-AI use, explicitly evaluating novelty/distinctiveness in review, and designing mechanisms to preserve exploratory projects.
- Caution on long-run impact
- Early publication gains at NIH are concentrated in lower-impact papers; there is no current evidence that LLM use raises the probability of producing highly cited, transformative work. Long-term effects on discovery and scientific impact remain uncertain.
- Research and policy priorities
- Need causal studies to separate writing-quality vs. substantive-idea mechanisms.
- Extend analyses to more universities, additional agencies, longer follow-up for citation outcomes, and evolving LLM architectures and detection methods.
- Consider experiments or pilot policies (disclosure, reviewer training, modified evaluation rubrics) to manage potential homogenization while preserving productivity benefits.
Limitations to bear in mind: LLM-detection is imperfect and model-dependent; private-proposal data come from two institutions (generalizability limit); observational design cannot establish causality; publication windows are short (1–2 years) for long-run impact assessment.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLM use rises sharply beginning in 2023. Adoption Rate | positive | LLM involvement / use over time |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLM use exhibits a bimodal distribution, indicating a clear split between minimal and substantive use. Adoption Rate | mixed | distributional pattern of LLM involvement (minimal vs. substantive) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency. Innovation Output | negative | semantic distinctiveness (distance from recently funded work) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| At NIH, LLM use is positively associated with proposal success. Research Productivity | positive | proposal success / funding probability |
Reading fidelity
high
Study strength
medium
|
not reported
|
| At NIH, LLM use is associated with higher early-stage publication output. Research Productivity | positive | early-stage publication output (number of publications soon after funding) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| No comparable positive associations between LLM use and proposal success or early-stage publication output are observed at NSF. Research Productivity | null_result | proposal success and early-stage publication output |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The productivity gains at NIH are concentrated in non-hit papers rather than the most highly cited work. Output Quality | mixed | citation impact distribution (non-hit vs highly cited papers) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The rise of LLMs is reshaping how scientific ideas are positioned, selected, and translated into publicly funded research, with implications for portfolio governance, research diversity, and the long-run impact of science. Governance And Regulation | mixed | direction/diversity/impact of scientific research portfolio (qualitative synthesis) |
Reading fidelity
high
Study strength
speculative
|
not reported
|