Large language models make scientific citations warmer and broader: in a position-matched audit of 1,746 recent NLP papers, LLMs replace many human critiques with endorsements, preferentially recommend older, high-impact work, and draw more on socially distant authors — amplifying visibility while diluting critical engagement.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.
Summary
Main Finding
Large language models (LLMs) systematically reshape scientific citation rhetoric and reach. When asked to reconstruct masked citation sentences in real paper contexts, LLMs (six tested) produce fewer critical/contrasting citations, favor older and more highly cited papers (amplifying visibility bias), and cite authors who are more socially distant in the coauthorship network than human authors do. These shifts are robust across models and conditioned by citation intent and paper section.
Key Points
-
Task and scale
- Position-aligned, slot-level masked-citation reconstruction on 1,746 ACL/EMNLP/NAACL 2025 main-track papers (.tex + .bib).
- Corpus: 63,944 citation contexts → 132,913 citation slots.
- Matched to Dimensions bibliometric records (human references matched 86.7%).
- Six LLMs evaluated: GPT-5.1, Claude-3.5-Haiku, Gemini-2.0-Flash, DeepSeek-V3.2, Llama-4-Maverick, Qwen2.5-72B-Instruct.
-
Rhetorical shift: "warming" of citation tone (RQ1)
- Human baseline intent distribution: ~21% supporting, 19% contrasting, 60% mentioning.
- LLMs produce far fewer contrasting citations (model range ≈ 10.7%–17.2%) and typically more supporting citations (several models produce supporting shares ≈ 38%–41%).
- Contrasting human citations are preserved by LLMs only ~34.6%–50.6% of the time; asymmetric moves favor contrasting→mentioning/supporting (12.5%–31.1% contrasting→supporting), while supporting→contrasting is rare (3.6%–7.6%).
- LLMs’ self-reported generation intent is even “warmer” (many reporting 54%–75% supporting).
-
Bias amplification by intent and paper attributes (RQ2)
- LLMs over-select papers that are older and have higher citation counts (visibility/popularity bias), compared with the human-chosen citations for the same slots.
- This popularity/recency amplification is especially pronounced when human authors use contrasting citations (humans tend to contrast against recent, niche work; LLMs substitute older, more prominent works).
-
Social reach: diminished local homophily (RQ3)
- Authors cited by humans are more likely to be socially close (coauthorship proximity) to the citing author—especially for supporting citations.
- LLM-generated citations tend to pick more socially distant authors in a 20.3M-edge coauthorship network of ~2.1M researchers, i.e., LLMs broaden reach beyond a citing author’s immediate collaborative neighborhood.
-
Methodological strengths
- Slot-level counterfactuals: each LLM citation directly replaces a known human choice in the same rhetorical context.
- Post-cutoff evaluation: chosen papers were released after the models’ training cutoffs to avoid memorized retrieval.
- Intent labeling via an independent LLM-as-judge, validated against human annotators (judge vs human majority κ ≈ 0.60 for supporting/contrasting).
Data & Methods
- Corpus: 1,746 main-track papers from ACL/EMNLP/NAACL 2025 with accessible source and bib files.
- Masked-citation task: for each citation sentence, LLMs receive surrounding one-sentence context, section heading, citing paper metadata, and required citation count; they must generate replacement citation sentences and list the referenced works.
- Models tested: GPT-5.1, Claude-3.5-Haiku, Gemini-2.0-Flash, DeepSeek-V3.2, Llama-4-Maverick, Qwen2.5-72B-Instruct. Citation-intent judge used Gemini-3-Flash-Preview (robustness check with DeepSeek-V4-Flash).
- Matching: references mapped to Dimensions by DOI or title to obtain metadata (year, team size, citation counts). Unmatched references excluded from downstream analyses.
- Coauthorship analysis: placed cited and citing authors into a 20.3M-edge coauthorship network spanning ~2.1M researchers to measure social distance (dyadic distances).
- Key validations: human annotation checks for intent labeling; multiple judges/replications to ensure robustness.
Implications for AI Economics
-
Redistribution of attention and cumulative advantage
- Amplified popularity bias means LLM-assisted writing can direct more citation attention to already-visible, older papers, strengthening Matthew-effect dynamics and potentially increasing concentration of academic attention and credit toward incumbents.
- This can affect economic allocation in academia (hiring, funding, promotions) that rely on citation-based indicators, potentially inflating returns to established work and slowing recognition of novel/niche contributions.
-
Incentives and signaling
- Reduced critical (contrasting) citations may dampen public scientific debate signaled in the literature. If LLMs produce fewer critical engagements, the visible record of scientific disagreement—important for signaling robustness and directing follow-on work—could shrink, altering incentives for replication and critical follow-up.
- The rhetorical warming may change how papers are perceived by readers and evaluators, affecting downstream demand for particular methods or topics.
-
Knowledge diffusion and market structure
- Broader social reach (LLMs citing more socially distant authors) could partially counterbalance concentration by exposing readers to authors outside a citing author’s immediate network—potentially improving cross-network diffusion and discovery.
- However, if reach broadens toward already-popular distant authors (rather than niche distant authors), the net effect may still favor centralization.
-
Policy, measurement, and tooling implications
- Bibliometric measures and funding/resource allocation formulas may need adjustment or robustness checks to account for increased automated citation generation and its biases.
- Transparency requirements (declaring LLM assistance in drafting citations) and provenance tooling (flagging which citations were suggested by models) can help evaluators interpret citation signals correctly.
- Practical mitigations: develop citation-assistant models/prompts that preserve rhetorical intent, prioritize recency/diversity heuristics, or surface niche/recent work and social-proximity signals to counteract popularity amplification.
- Auditing and regulation: ongoing audits of LLM effects on bibliometrics should be part of research-evaluation governance to avoid unintended market distortions.
In short: LLMs change both the tone and the destination of citations—making scholarly claims look more supported and steering attention toward older, popular works while also reaching further across social networks. These shifts carry material economic consequences for the distribution of credit, attention, and resources in science and argue for transparency, tooling, and policy responses.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Compared with human-written citations, all six tested LLMs produce fewer contrasting citations. Research Productivity | negative | Share of citations classified as contrasting or critical |
Reading fidelity
high
Study strength
medium
|
n=1746
10.7–17.2% contrasting citations for LLMs versus 19.3% for humans
|
| Five of the six tested LLMs produce more supporting citations than the human-written baseline. Research Productivity | positive | Share of citations classified as supporting |
Reading fidelity
high
Study strength
medium
|
n=1746
Five of six models above the human 21.3% baseline, up to 40.6%
|
| LLM-generated citation sentences exhibit a persistent directional warming of rhetorical intent relative to human-written sentences. Research Productivity | negative | Change in citation rhetorical intent from contrasting toward mentioning or supporting |
Reading fidelity
high
Study strength
medium
|
n=132913
∆cont = −2.4 to −8.6%
|
| Human contrasting citations are more often converted into supporting citations than human supporting citations are converted into contrasting citations. Research Productivity | negative | Asymmetric transitions between contrasting and supporting citation intent |
Reading fidelity
high
Study strength
medium
|
n=132913
Contrasting→supporting: 12.5–31.1%; supporting→contrasting: 3.6–7.6%
|
| Contrasting citation intent is less likely to be preserved by LLMs than mentioning intent. Research Productivity | negative | Intent-label preservation between human and LLM-generated citation sentences |
Reading fidelity
high
Study strength
medium
|
n=132913
Contrasting preservation: 34.6–50.6%; mentioning preservation: 60–80%
|
| LLM-generated citations have lower bibliometric matching rates than human citations, indicating more hallucinated or malformed references. Error Rate | negative | Rate of generated citations matched to a bibliometric database record |
Reading fidelity
high
Study strength
medium
|
n=132913
Human: 86.7% matched; LLMs: 39.5–81.9% matched
|
| LLMs preferentially cite older and more popular papers, and this tendency is amplified for contrasting citations. Inequality | negative | Age and citation popularity of selected cited papers |
Reading fidelity
high
Study strength
medium
|
n=132913
|
| Humans tend to cite papers within their close coauthorship-based social networks, especially for supporting citations, whereas LLMs tend to cite authors who are more socially distant. Research Productivity | mixed | Coauthorship-network social distance between citing and cited authors |
Reading fidelity
high
Study strength
medium
|
n=2100000
|
| The citation differences observed between humans and LLMs are systematic across models and depend on citation intent. Research Productivity | mixed | Human–LLM divergence in citation selection and rhetorical intent conditional on citation intent |
Reading fidelity
high
Study strength
medium
|
n=1746
|