0 cumulative citations
View corpus contextGenerative AI is boosting individual scientific output and citations but risks homogenising research, degrading training, and undermining selection signals; the authors propose a four-principle Responsible Research with AI framework to preserve gains while managing systemic risks.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper examines the tension between the benefits of generative artificial intelligence (AI) for scientific research and the unresolved governance questions that accompany its rapid adoption. Drawing on an academic roundtable held at the AI for Science and Innovation Workshop (Scuola IMT Alti Studi Lucca, April 2026) and on a fast-expanding empirical literature, it maps the disagreement within the research community across four stages of the research process: funding, research tasks, publication and peer review, and use and uptake. The empirical case for AI's productivity, augmentation, and democratization effects has strengthened. The picture changes once productivity is disaggregated: AI-assisted work shows measurable gains in publication volume and citation share, while the evidence on novelty, disruption, and breakthrough output remains ambiguous or negative. We argue that the divergence between private and social returns arises through three analytically distinct mechanisms, namely information asymmetry, negative externalities on a shared knowledge base, and depletion of research capacity, and that each calls for a different governance instrument. We propose Responsible Research with AI (RRAI), an extension of the Responsible Research and Innovation tradition organized around four principles that operate at different levels of the research system: disclosure, differentiation, narrative, and proportionality. RRAI builds on existing institutional scaffolding, including the EU AI Act, UNESCO, and the OECD, and aims to preserve AI's productivity gains while addressing systemic risks that individual researchers can neither observe nor manage on their own.
Summary
Main Finding
Generative AI delivers clear individual-level productivity gains (more papers, higher citation shares, faster writing/coding) but creates systemic risks that can reduce the social value of research. These risks operate through three distinct economic mechanisms — information asymmetry (adverse selection), negative externalities on the shared knowledge base (homogenization/model collapse), and depletion of research capacity (weakened training) — and therefore require different governance instruments. The authors propose Responsible Research with AI (RRAI), a four-principle framework (disclosure, differentiation, narrative, proportionality) that builds on existing institutions (EU AI Act, UNESCO, OECD) to preserve private gains while managing collective harms.
Key Points
- Empirical pattern (aggregate):
- AI-assisted researchers publish more and capture larger citation shares, but evidence on novelty, disruption, and breakthrough output is weak or negative.
- Examples: LLM adoption associated with publication increases of 23.7–89.3% (Kusumegi et al., 2025); AI-augmented researchers publish ~3.02× more papers and receive ~4.84× more citations, while topic volume shrinks 4.63% and cross-scientist engagement falls 22% (Hao et al., 2026).
- Disclosure of AI use is rare (e.g., ~1.5% in medical manuscripts), and detection is unreliable; fabricated references in audits rose to ~1 per 277 papers in early 2026 (audit cited).
- Three analytically distinct mechanisms creating divergence between private and social returns:
- Information asymmetry / adverse selection: undisclosed AI assistance undermines evaluators' ability to infer researchers' underlying capability from past output, biasing hiring/granting toward skilled AI users rather than skilled researchers.
- Negative externalities on the knowledge commons: AI-mediated homogenization and feedback into training corpora can reduce topical diversity and amplify correlated model errors (model collapse), degrading the scientific corpus for future work and model training.
- Capacity depletion (human capital): automation of the very tasks through which researchers acquire expertise (writing, debugging, hypothesis formation) can short-circuit learning unless engagement modes prioritize explanation and interrogation over copying.
- Self-regulation has been fast but ineffective: publishers and COPE issued disclosure policies quickly, yet compliance remains low because disclosure imposes private costs while benefits are collective — a commons problem not easily solved by exhortation.
- Workload and funding implications: AI reconfigures tasks (substitutes some, augments others, creates verification/compliance tasks), often increasing net effort and potentially worsening attribution problems for funders; this could lead to reduced public research budgets if outcomes appear less attributable to investment.
- RRAI framework: four principles targeted to mechanisms:
- Disclosure (informational level): mandatory, standardized, verifiable reporting of AI use to correct adverse selection.
- Differentiation (pedagogical level): training and norms that distinguish mode-of-engagement to preserve learning (address capacity depletion).
- Narrative (communicative level): public accounting and explanation of AI's role to align incentives and maintain trust (mitigate funding/attribution risks).
- Proportionality (distributional level): policy measures calibrated to the scale of systemic risk (target negative externalities and concentration effects).
Data & Methods
- Primary basis: synthesis of a fast-expanding empirical literature plus insights from an academic roundtable at the AI for Science and Innovation Workshop (Scuola IMT, April 2026). Roundtable consisted mainly of economists and innovation scholars; pre- and post-debate votes shifted slightly (10→9 in favor; 6→8 against).
- Evidence types surveyed and cited:
- Bibliometrics and large-audience audits (e.g., analyses of millions of preprints/papers; Hao et al., 2026; Kusumegi et al., 2025).
- Experimental studies (randomized/ preregistered experiments on LLM-assisted programming/writing — e.g., Lehmann et al., 2025; Noy & Zhang, 2023).
- Field and lab evidence on productivity and task substitution (Brynjolfsson et al., 2025; Cockburn et al., 2018).
- Detection/audit studies for undisclosed AI use and fabricated references (Lu et al., 2026; PubMed audit cited).
- Theoretical and simulation analyses (models of prioritized search, task-based models of productivity, and simulations of submission pipelines).
- Complementary literature from ethics, STS, and science policy (Resnik & Hosseini, 2025; Owen et al., 2012; Stilgoe et al., 2013).
- Limitations noted by authors:
- Heterogeneity across fields, tasks, and modes of AI engagement; many empirical studies are recent and evolving.
- Detection methods vary widely and benchmarks are unsettled.
- Roundtable is informative but not representative; many empirical claims are proxied from early studies and preprints.
Implications for AI Economics
- Reconceptualize productivity gains: Economists should distinguish between private marginal product of AI use (individual output, citations) and social marginal product (novelty, robustness, diversity). Standard productivity metrics (publications, citations) may overstate social returns when AI widens the production-progress gap.
- Information asymmetry and labor market signaling:
- Undisclosed AI use creates adverse selection in hiring/grants; economists should model how disclosure rules or detection technologies change equilibrium wages, promotion, and allocation.
- Design of credible signals (standardized AI-use reporting, verifiable provenance) becomes an economic policy instrument to restore selection efficiency.
- Knowledge commons externalities:
- Model the dynamic feedback loop between published AI-assisted work, training corpora for future models, and homogenization/model collapse. Quantify welfare losses from reduced topical diversity and correlated errors.
- Consider antitrust-like or data-governance interventions that preserve diversity of training corpora and prevent monoculture effects.
- Human capital formation:
- Study how different modes of AI engagement (explanatory vs. generative copying) affect skill accumulation and long-term researcher productivity. This calls for dynamic models of human capital with task substitution that allow for skill depreciation or stunting.
- Evaluate pedagogical interventions (differentiated training, supervised-use requirements) as policy tools; estimate long-run returns to academic labor markets.
- Public funding and attribution:
- AI complicates attribution of outcomes to public research spending. Economists should develop evaluation metrics and accountability mechanisms that reduce the risk of underinvestment due to misattribution.
- Explore funding reforms: conditional grants tied to verifiable methodological disclosure, support for high-risk/high-novelty research less amenable to AI shortcuts, and incentives to maintain topical diversity.
- Regulatory design and instrument matching:
- Match instruments to mechanisms: mandatory disclosure and reporting standards to correct information asymmetry; pedagogical/differentiation policies (curriculum, doctoral supervision standards) to address capacity depletion; public-good-oriented data and model governance (limits on dataset reuse or enforced diversity) to reduce negative externalities; proportional regulatory measures that scale with systemic risk.
- Leverage existing supranational frameworks (EU AI Act, UNESCO, OECD) but adapt them to the research context (e.g., standardized AI-use metadata for manuscripts, audits, and model training datasets).
- Research agenda suggestions:
- Empirically estimate social vs private returns to AI in research across fields, using difference-in-differences and selection-corrected designs.
- Measure prevalence and modes of undisclosed AI use with robust detection benchmarks and field surveys.
- Quantify topic concentration dynamics and trainset contamination effects on future model performance.
- Evaluate interventions (mandatory disclosure, pedagogical guidance, funding rule changes) with randomized or quasi-experimental designs.
- Model equilibrium effects on academic labor markets and research funding when AI alters signaling and productivity measures.
Summary takeaway: Generative AI can accelerate individual research productivity but risks systemic harms through distinct economic mechanisms. Policy should not be one-size-fits-all — instruments must be matched to mechanism (disclosure for adverse selection, pedagogical differentiation for capacity depletion, data/model governance for knowledge-commons externalities) to preserve social returns while enabling continued individual gains.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AI-assisted research is associated with increases in publication output ranging from 23.7% to 89.3%, depending on scientific field and author background. Research Productivity | positive | Publication output per researcher |
Reading fidelity
high
Study strength
high
|
n=2100000
23.7 to 89.3 percent
|
| Researchers engaged in AI-augmented work publish 3.02 times more papers. Research Productivity | positive | Number of papers published |
Reading fidelity
high
Study strength
high
|
n=41300000
3.02 times more papers
|
| Researchers engaged in AI-augmented work receive 4.84 times more citations. Research Productivity | positive | Citations received |
Reading fidelity
high
Study strength
high
|
n=41300000
4.84 times more citations
|
| AI adoption reduces the collective volume of scientific topics studied by 4.63%. Research Productivity | negative | Collective diversity or volume of scientific topics studied |
Reading fidelity
high
Study strength
high
|
n=41300000
4.63 percent reduction
|
| AI adoption reduces engagement between scientists by 22%. Team Performance | negative | Engagement or interaction between scientists |
Reading fidelity
high
Study strength
high
|
n=41300000
22 percent reduction
|
| AI-assisted research has not produced analogous gains in novelty or disruption, and breakthrough output has not accelerated commensurately with publication volume. Innovation Output | null_result | Research novelty, disruption, and breakthrough output |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Only 1.5% of medical manuscripts disclose LLM involvement. Governance And Regulation | negative | Disclosure of LLM involvement in manuscripts |
Reading fidelity
high
Study strength
medium
|
1.5 percent
|
| Revealed AI assistance reduces perceived manuscript quality by 0.32 points on a five-point scale. Output Quality | negative | Perceived manuscript quality |
Reading fidelity
high
Study strength
medium
|
0.32 points on a five-point scale
|
| An audit found one fabricated reference per 277 PubMed-indexed papers in the first weeks of 2026, compared with one per 2,828 papers in 2023. Error Rate | negative | Rate of fabricated references |
Reading fidelity
high
Study strength
medium
|
n=2500000
one fabricated reference per 277 papers, up from one per 2,828 in 2023
|
| In two pre-registered programming-learning experiments, there was no difference in average learning outcomes between subjects who learned with and without access to an LLM. Skill Acquisition | null_result | Average programming learning outcomes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Within the programming-learning experiments, students who used the model to generate solutions covered more material while understanding less, whereas students who used it to request explanations deepened their understanding. Skill Acquisition | mixed | Material covered and depth of understanding |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In a small neuroimaging study of essay writing, participants using an LLM showed weaker brain connectivity, lower recall of their own text, and a reduced sense of ownership, with effects persisting after the tool was withdrawn. Skill Obsolescence | negative | Brain connectivity, recall of written text, and sense of ownership |
Reading fidelity
high
Study strength
low
|
not reported
|
| In a survey of researchers, 85% wanted regulatory guidance, 58.7% wanted guidance specifically from supranational institutions, and 70% identified legal uncertainty as the principal barrier to adoption. Governance And Regulation | mixed | Demand for AI regulation and perceived barriers to AI adoption |
Reading fidelity
high
Study strength
medium
|
n=6215
85 percent; 58.7 percent; 70 percent
|