The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models make scientific citations warmer and broader: in a position-matched audit of 1,746 recent NLP papers, LLMs replace many human critiques with endorsements, preferentially recommend older, high-impact work, and draw more on socially distant authors — amplifying visibility while diluting critical engagement.

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Yixuan Liu, Lin Chen, Zhuoqi Liu, Jianglin Lu, Dakota Murray · September 01, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yixuan Liu unresolved corpus identity
  2. Lin Chen unresolved corpus identity
  3. Zhuoqi Liu unresolved corpus identity
  4. Jianglin Lu unresolved corpus identity
  5. Dakota Murray unresolved corpus identity
Using a slot-level counterfactual reconstruction across 1,746 NLP conference papers and six LLMs, the authors find that LLM-generated citation sentences are systematically less critical (fewer contrasting citations), overrepresent older and popular papers, and cite authors who are more socially distant in the coauthorship network.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

Summary

Main Finding

Large language models (LLMs) systematically reshape scientific citation rhetoric and reach. When asked to reconstruct masked citation sentences in real paper contexts, LLMs (six tested) produce fewer critical/contrasting citations, favor older and more highly cited papers (amplifying visibility bias), and cite authors who are more socially distant in the coauthorship network than human authors do. These shifts are robust across models and conditioned by citation intent and paper section.

Key Points

  • Task and scale

    • Position-aligned, slot-level masked-citation reconstruction on 1,746 ACL/EMNLP/NAACL 2025 main-track papers (.tex + .bib).
    • Corpus: 63,944 citation contexts → 132,913 citation slots.
    • Matched to Dimensions bibliometric records (human references matched 86.7%).
    • Six LLMs evaluated: GPT-5.1, Claude-3.5-Haiku, Gemini-2.0-Flash, DeepSeek-V3.2, Llama-4-Maverick, Qwen2.5-72B-Instruct.
  • Rhetorical shift: "warming" of citation tone (RQ1)

    • Human baseline intent distribution: ~21% supporting, 19% contrasting, 60% mentioning.
    • LLMs produce far fewer contrasting citations (model range ≈ 10.7%–17.2%) and typically more supporting citations (several models produce supporting shares ≈ 38%–41%).
    • Contrasting human citations are preserved by LLMs only ~34.6%–50.6% of the time; asymmetric moves favor contrasting→mentioning/supporting (12.5%–31.1% contrasting→supporting), while supporting→contrasting is rare (3.6%–7.6%).
    • LLMs’ self-reported generation intent is even “warmer” (many reporting 54%–75% supporting).
  • Bias amplification by intent and paper attributes (RQ2)

    • LLMs over-select papers that are older and have higher citation counts (visibility/popularity bias), compared with the human-chosen citations for the same slots.
    • This popularity/recency amplification is especially pronounced when human authors use contrasting citations (humans tend to contrast against recent, niche work; LLMs substitute older, more prominent works).
  • Social reach: diminished local homophily (RQ3)

    • Authors cited by humans are more likely to be socially close (coauthorship proximity) to the citing author—especially for supporting citations.
    • LLM-generated citations tend to pick more socially distant authors in a 20.3M-edge coauthorship network of ~2.1M researchers, i.e., LLMs broaden reach beyond a citing author’s immediate collaborative neighborhood.
  • Methodological strengths

    • Slot-level counterfactuals: each LLM citation directly replaces a known human choice in the same rhetorical context.
    • Post-cutoff evaluation: chosen papers were released after the models’ training cutoffs to avoid memorized retrieval.
    • Intent labeling via an independent LLM-as-judge, validated against human annotators (judge vs human majority κ ≈ 0.60 for supporting/contrasting).

Data & Methods

  • Corpus: 1,746 main-track papers from ACL/EMNLP/NAACL 2025 with accessible source and bib files.
  • Masked-citation task: for each citation sentence, LLMs receive surrounding one-sentence context, section heading, citing paper metadata, and required citation count; they must generate replacement citation sentences and list the referenced works.
  • Models tested: GPT-5.1, Claude-3.5-Haiku, Gemini-2.0-Flash, DeepSeek-V3.2, Llama-4-Maverick, Qwen2.5-72B-Instruct. Citation-intent judge used Gemini-3-Flash-Preview (robustness check with DeepSeek-V4-Flash).
  • Matching: references mapped to Dimensions by DOI or title to obtain metadata (year, team size, citation counts). Unmatched references excluded from downstream analyses.
  • Coauthorship analysis: placed cited and citing authors into a 20.3M-edge coauthorship network spanning ~2.1M researchers to measure social distance (dyadic distances).
  • Key validations: human annotation checks for intent labeling; multiple judges/replications to ensure robustness.

Implications for AI Economics

  • Redistribution of attention and cumulative advantage

    • Amplified popularity bias means LLM-assisted writing can direct more citation attention to already-visible, older papers, strengthening Matthew-effect dynamics and potentially increasing concentration of academic attention and credit toward incumbents.
    • This can affect economic allocation in academia (hiring, funding, promotions) that rely on citation-based indicators, potentially inflating returns to established work and slowing recognition of novel/niche contributions.
  • Incentives and signaling

    • Reduced critical (contrasting) citations may dampen public scientific debate signaled in the literature. If LLMs produce fewer critical engagements, the visible record of scientific disagreement—important for signaling robustness and directing follow-on work—could shrink, altering incentives for replication and critical follow-up.
    • The rhetorical warming may change how papers are perceived by readers and evaluators, affecting downstream demand for particular methods or topics.
  • Knowledge diffusion and market structure

    • Broader social reach (LLMs citing more socially distant authors) could partially counterbalance concentration by exposing readers to authors outside a citing author’s immediate network—potentially improving cross-network diffusion and discovery.
    • However, if reach broadens toward already-popular distant authors (rather than niche distant authors), the net effect may still favor centralization.
  • Policy, measurement, and tooling implications

    • Bibliometric measures and funding/resource allocation formulas may need adjustment or robustness checks to account for increased automated citation generation and its biases.
    • Transparency requirements (declaring LLM assistance in drafting citations) and provenance tooling (flagging which citations were suggested by models) can help evaluators interpret citation signals correctly.
    • Practical mitigations: develop citation-assistant models/prompts that preserve rhetorical intent, prioritize recency/diversity heuristics, or surface niche/recent work and social-proximity signals to counteract popularity amplification.
    • Auditing and regulation: ongoing audits of LLM effects on bibliometrics should be part of research-evaluation governance to avoid unintended market distortions.

In short: LLMs change both the tone and the destination of citations—making scholarly claims look more supported and steering attention toward older, popular works while also reaching further across social networks. These shifts carry material economic consequences for the distribution of credit, attention, and resources in science and argue for transparency, tooling, and policy responses.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large, systematic dataset and several robustness checks (multiple LLMs, post-cutoff evaluation, count- and position-controlled reconstruction, judge validation against human annotators, coauthorship network), which yield consistent patterns; however, causal claims about downstream effects are limited, results rely on LLM-as-judge labeling (which can introduce bias), matching rates vary across models (hallucinations), and the sample is restricted to NLP conference papers, limiting external validity. Methods Rigorhigh — Carefully designed counterfactual reconstruction with position- and count-matching and post-cutoff papers prevents direct retrieval; multiple LLMs and judges used with human annotation validation; rich bibliometric grounding and a very large coauthorship network allow structural analyses. Remaining issues include reliance on LLM judges for intent labeling, variable model matching rates (hallucinations), potential sensitivity to prompt design and hyperparameters, and domain restriction to NLP conference papers. Sample1,746 main-track papers from ACL, EMNLP, and NAACL 2025 with .tex and .bib from arXiv, yielding 63,944 citation contexts and 132,913 citation slots; 86.7% of human citations matched to Dimensions records; six LLMs evaluated (GPT-5.1, Claude-3.5-Haiku, Gemini-2.0-Flash, DeepSeek-V3.2, Llama-4-Maverick, Qwen2.5-72B-Instruct); citation intent labeled by an independent LLM judge (Gemini-3-Flash-Preview with robustness checks) and validated against human annotators; matched references linked to Dimensions metadata and placed in a 20.3M-edge coauthorship network (2.1M researchers) for social-distance measures. Themesadoption governance IdentificationPosition-aligned masked-citation reconstruction: for each citation sentence in a held-out corpus, the citation sentence is masked and multiple LLMs are prompted to reconstruct the same-number-of-citations sentence (post-training-cutoff), producing counterfactual LLM-written citations directly comparable to the original human citation; an independent LLM judge (and human annotators for validation) labels rhetorical intent; matched references are linked to Dimensions for metadata and authors are placed in a 20.3M-edge coauthorship network to measure social distance. GeneralizabilityRestricted to NLP conference papers (ACL/EMNLP/NAACL 2025) — may not generalize to other disciplines, publication types, or less-formal venues., English-language, main-track accepted papers only; excludes unpublished manuscripts and diverse writing conventions., LLM behavior may vary with prompting, temperature, model versions, access to tools or retrieval augmentation — results depend on the specific prompts and models tested., Intent labeling relies primarily on an LLM-as-judge (with some human validation), introducing possible systematic labeling biases., Bibliometric matching depends on Dimensions coverage and DOI/title matching; hallucinated or malformed LLM references reduce match rates and could bias downstream analyses., Post-cutoff design prevents direct memorization but may not reflect current/future models trained on these papers.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Compared with human-written citations, all six tested LLMs produce fewer contrasting citations. Research Productivity negative Share of citations classified as contrasting or critical
Reading fidelity high
Study strength medium
n=1746
10.7–17.2% contrasting citations for LLMs versus 19.3% for humans
0.18
Five of the six tested LLMs produce more supporting citations than the human-written baseline. Research Productivity positive Share of citations classified as supporting
Reading fidelity high
Study strength medium
n=1746
Five of six models above the human 21.3% baseline, up to 40.6%
0.18
LLM-generated citation sentences exhibit a persistent directional warming of rhetorical intent relative to human-written sentences. Research Productivity negative Change in citation rhetorical intent from contrasting toward mentioning or supporting
Reading fidelity high
Study strength medium
n=132913
∆cont = −2.4 to −8.6%
0.18
Human contrasting citations are more often converted into supporting citations than human supporting citations are converted into contrasting citations. Research Productivity negative Asymmetric transitions between contrasting and supporting citation intent
Reading fidelity high
Study strength medium
n=132913
Contrasting→supporting: 12.5–31.1%; supporting→contrasting: 3.6–7.6%
0.18
Contrasting citation intent is less likely to be preserved by LLMs than mentioning intent. Research Productivity negative Intent-label preservation between human and LLM-generated citation sentences
Reading fidelity high
Study strength medium
n=132913
Contrasting preservation: 34.6–50.6%; mentioning preservation: 60–80%
0.18
LLM-generated citations have lower bibliometric matching rates than human citations, indicating more hallucinated or malformed references. Error Rate negative Rate of generated citations matched to a bibliometric database record
Reading fidelity high
Study strength medium
n=132913
Human: 86.7% matched; LLMs: 39.5–81.9% matched
0.18
LLMs preferentially cite older and more popular papers, and this tendency is amplified for contrasting citations. Inequality negative Age and citation popularity of selected cited papers
Reading fidelity high
Study strength medium
n=132913
0.18
Humans tend to cite papers within their close coauthorship-based social networks, especially for supporting citations, whereas LLMs tend to cite authors who are more socially distant. Research Productivity mixed Coauthorship-network social distance between citing and cited authors
Reading fidelity high
Study strength medium
n=2100000
0.18
The citation differences observed between humans and LLMs are systematic across models and depend on citation intent. Research Productivity mixed Human–LLM divergence in citation selection and rhetorical intent conditional on citation intent
Reading fidelity high
Study strength medium
n=1746
0.18

Notes