The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

The spread of large language models is changing how grant ideas are framed and selected: proposers who use LLMs produce less semantically distinctive submissions and—at NIH but not NSF—are more likely to win funding and generate early, modest-impact papers.

The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
Yifan Qian, Zhe Wen, Alexander C. Furnas, Yue Bai, Erzhuo Shao, Dashun Wang · January 21, 2026
arxiv correlational medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yifan Qian unresolved corpus identity
  2. Zhe Wen unresolved corpus identity
  3. Alexander C. Furnas unresolved corpus identity
  4. Yue Bai unresolved corpus identity
  5. Erzhuo Shao unresolved corpus identity
  6. Dashun Wang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Yifan Qian provider ID
  2. Zheng Wen provider ID
  3. Alexander C. Furnas provider ID
  4. Yu Bai provider ID
  5. Erzhuo Shao provider ID
  6. Dashun Wang provider ID
Rising LLM use in grant-writing since 2023 is linked to lower semantic distinctiveness of proposals and, depending on agency, to higher award rates and early publication output (positive at NIH but not NSF), with NIH gains concentrated in moderately cited work rather than top hits.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly diffusing into scientific practice, holding substantial promise while raising widespread concerns. Despite growing attention to AI use in scientific writing and evaluation, little is known about how the rise of LLMs is reshaping the public funding landscape. Here, we examine LLM involvement at key stages of the federal funding pipeline by combining two complementary data sources: confidential National Science Foundation (NSF) and National Institutes of Health (NIH) proposal submissions from two large US R1 universities, including funded, unfunded, and pending proposals, and the full population of publicly released NSF and NIH awards. We find that LLM use rises sharply beginning in 2023 and exhibits a bimodal distribution, indicating a clear split between minimal and substantive use. Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency. The consequences of this shift are agency-dependent. LLM use is positively associated with proposal success and higher early-stage publication output at NIH, whereas no comparable associations are observed at NSF. Notably, the productivity gains at NIH are concentrated in non-hit papers rather than the most highly cited work. Together, these findings provide large-scale evidence that the rise of LLMs is reshaping how scientific ideas are positioned, selected, and translated into publicly funded research, with implications for portfolio governance, research diversity, and the long-run impact of science.

Summary

Main Finding

The rapid adoption of large language models (LLMs) in federal grant writing since 2023 is detectable and uneven (bimodal). Higher LLM involvement is consistently associated with proposals and awards that are semantically closer to recently funded work (lower distinctiveness). Consequences differ by agency: at NIH higher LLM involvement is associated with greater proposal success (~+4 percentage points) and ~+5% more early-stage publications (driven by non-hit papers), whereas at NSF no systematic association with funding success or publication output is observed.

Key Points

  • Temporal pattern
    • Detectable LLM use rose sharply beginning in 2023 (coinciding with ChatGPT availability).
    • Individual-grant LLM involvement (α) shows a bimodal distribution in 2023–2025: one mode near zero, another around ~10–15% α.
  • Positioning of ideas
    • Greater LLM involvement is associated with lower semantic distinctiveness (proposals/awards move closer to the agency’s recent funded portfolio).
    • Magnitude: moving an award from the 25th to the 75th percentile of α corresponds to ~5 percentage-point decrease in distinctiveness for NSF and ~4 points for NIH (within-year percentile rank).
    • Rewriting 2021 abstracts with an LLM did not materially change distinctiveness measures, suggesting substantive convergence—not only surface-level stylistic smoothing.
  • Selection and outputs (agency-dependent)
    • NSF: no statistically significant relationship between α and proposal success or publication counts.
    • NIH: higher α is associated with higher funding probability (~+4 percentage points from 25th→75th α) and ~+5% more early publications for grants (2023–2024 cohorts).
    • The NIH publication gain is concentrated in non-hit papers; associations with top 5% (or top 1%) high-impact papers are not significant.
  • Robustness
    • Results robust to controls (grant-year, field, investigator fixed effects; funding amount).
    • LLM-associated proposals score as more readable (Flesch reading ease), and removing promotional words does not change the LLM-use signal.

Data & Methods

  • Data
    • Private submissions: complete NSF and NIH proposal abstracts (funded, unfunded, pending) from two large U.S. R1 universities (request start dates 2021–2025).
    • Public awards: full population of publicly released NSF and NIH award abstracts (same period) plus linked publications for awards.
  • Measuring LLM involvement (α)
    • Based on Liang et al. detection approach: estimate word distributions for human-written text (public 2021 abstracts) and for LLM-generated text (GPT-3.5-turbo-0125 rewriting of the same abstracts).
    • Applied detection at corpus level (3-month rolling windows) and per-grant level to get α (fraction of LLM-modified sentences).
  • Semantic distinctiveness
    • Embed abstracts using SPECTER2 embeddings and compute average cosine distance to all agency-funded abstracts in the prior year.
    • Convert distances to within-year percentile ranks (higher = more distinctive).
  • Statistical analysis
    • Regressions of distinctiveness, funding success, and publication outcomes on grant-level α.
    • All regressions include start-year, field, and investigator (PI/co-PI) fixed effects and controls for requested/awarded funding amount.
    • Analyses exploit within-investigator variation in α (estimates are conditional correlations, not causal effects).
  • Robustness checks
    • Rewrote 2021 abstracts with LLMs to test whether surface-level changes explain distinctiveness shifts.
    • Evaluated reading-ease (Flesch) and removed promotional language; results held.

Implications for AI Economics

  • Upstream effects of LLMs matter: LLMs are already reshaping the funding pipeline (positioning and selection), not just downstream manuscript production.
  • Tradeoff: individual-level gains vs. system-level diversity
    • LLMs can raise submission quality/readability and may increase funded outputs (NIH), but they also appear to pull proposals toward established funded norms, reducing semantic novelty. This suggests a shift from exploration toward exploitation in public research portfolios.
  • Agency heterogeneity matters for outcomes and policy
    • Differences between NIH and NSF outcomes indicate that institutional review processes, norms, and evaluation criteria interact with LLM effects. Policy responses likely need to be agency-specific.
  • Implications for portfolio governance and evaluation
    • Funders should monitor LLM use and its impact on portfolio diversity and long-run innovation. Consider requiring disclosure of generative-AI use, explicitly evaluating novelty/distinctiveness in review, and designing mechanisms to preserve exploratory projects.
  • Caution on long-run impact
    • Early publication gains at NIH are concentrated in lower-impact papers; there is no current evidence that LLM use raises the probability of producing highly cited, transformative work. Long-term effects on discovery and scientific impact remain uncertain.
  • Research and policy priorities
    • Need causal studies to separate writing-quality vs. substantive-idea mechanisms.
    • Extend analyses to more universities, additional agencies, longer follow-up for citation outcomes, and evolving LLM architectures and detection methods.
    • Consider experiments or pilot policies (disclosure, reviewer training, modified evaluation rubrics) to manage potential homogenization while preserving productivity benefits.

Limitations to bear in mind: LLM-detection is imperfect and model-dependent; private-proposal data come from two institutions (generalizability limit); observational design cannot establish causality; publication windows are short (1–2 years) for long-run impact assessment.

Assessment

Paper Typecorrelational Evidence Strengthmedium — Large-scale, novel datasets (confidential submissions plus full public awards) provide strong coverage and clear, consistent associations across multiple outcomes; however, the design is observational without exogenous variation, leaving open confounding, selection, measurement error in LLM use, and reverse causality—so evidence supports robust associations but not definitive causal claims. Methods Rigormedium — The study leverages complementary data sources, constructs quantitative measures (LLM involvement and semantic distinctiveness), and compares patterns across agencies (NIH vs NSF) and outcomes (award rates, early publications, citation tiers), suggesting careful empirical work and robustness checks; nevertheless, key identification challenges (unobserved confounders, self-selection into LLM use, potential misclassification of LLM involvement) limit the overall rigor for causal inference. SampleConfidential proposal submissions to NSF and NIH (including funded, unfunded, and pending applications) from two large U.S. R1 universities, combined with the full population of publicly released NSF and NIH awards; temporal coverage spans the period before and after a sharp rise in LLM use beginning in 2023 (exact years and sample sizes not specified in the summary). Themesinnovation governance adoption productivity human_ai_collab IdentificationObservational comparison using two complementary datasets: confidential NSF and NIH proposal submissions (funded, unfunded, pending) from two large US R1 universities plus the full population of publicly released NSF and NIH awards; LLM involvement measured from proposal/award text and summarized (bimodal distribution); statistical associations estimated via regressions (controls for agency, field/discipline, time trends, and observable proposal characteristics) and cross-agency comparisons to probe robustness. No randomized assignment, instrumental variable, or other strong exogenous source of variation is reported, so causal claims rely on controlling for observables and comparative patterns across agencies. GeneralizabilityConfidential submission sample limited to two R1 universities—may not represent smaller institutions, non-US institutions, or broader researcher populations, Findings are specific to NSF and NIH funding ecosystems and may not generalize to other funders, countries, or private-sector R&D, Agency-specific effects (NIH vs NSF) imply discipline-specific differences (biomedical vs other sciences) that limit cross-field generalization, Measurement of LLM involvement from proposal text may miss private or unreported use and varies with disclosure norms, Temporal dynamics (early adoption period around 2023) may not reflect long-run equilibria as practices and policies evolve

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLM use rises sharply beginning in 2023. Adoption Rate positive LLM involvement / use over time
Reading fidelity high
Study strength medium
not reported
0.3
LLM use exhibits a bimodal distribution, indicating a clear split between minimal and substantive use. Adoption Rate mixed distributional pattern of LLM involvement (minimal vs. substantive)
Reading fidelity high
Study strength medium
not reported
0.3
Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency. Innovation Output negative semantic distinctiveness (distance from recently funded work)
Reading fidelity high
Study strength medium
not reported
0.3
At NIH, LLM use is positively associated with proposal success. Research Productivity positive proposal success / funding probability
Reading fidelity high
Study strength medium
not reported
0.3
At NIH, LLM use is associated with higher early-stage publication output. Research Productivity positive early-stage publication output (number of publications soon after funding)
Reading fidelity high
Study strength medium
not reported
0.3
No comparable positive associations between LLM use and proposal success or early-stage publication output are observed at NSF. Research Productivity null_result proposal success and early-stage publication output
Reading fidelity high
Study strength medium
not reported
0.3
The productivity gains at NIH are concentrated in non-hit papers rather than the most highly cited work. Output Quality mixed citation impact distribution (non-hit vs highly cited papers)
Reading fidelity high
Study strength medium
not reported
0.3
The rise of LLMs is reshaping how scientific ideas are positioned, selected, and translated into publicly funded research, with implications for portfolio governance, research diversity, and the long-run impact of science. Governance And Regulation mixed direction/diversity/impact of scientific research portfolio (qualitative synthesis)
Reading fidelity high
Study strength speculative
not reported
0.05

Notes