The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Detected LLM use on arXiv associates with a sustained rise in authors' paper output that exceeds what a stopping-time timing artifact would produce; calibrated placebos and several conservative designs preserve the positive association though causality is not claimed.

A robust association between LLM use and scientific productivity: Assessing stopping-time selection
Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin · July 31, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Keigo Kusumegi unresolved corpus identity
  2. Xinyu Yang unresolved corpus identity
  3. Paul Ginsparg unresolved corpus identity
  4. Mathijs de Vaan unresolved corpus identity
  5. Toby Stuart unresolved corpus identity
  6. Yian Yin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Keigo Kusumegi provider ID
  2. Xinyu Yang provider ID
  3. P. Ginsparg provider ID
  4. M. D. Vaan provider ID
  5. Toby Stuart provider ID
  6. Yian Yin provider ID
Using multiple robustness checks, calibrated placebos, and alternative estimators on arXiv data, the authors find a persistent positive association between detected LLM use and author productivity that cannot be explained solely by the stopping-time selection artifact identified by RBB.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.

Summary

Main Finding

Kusumegi et al. (2026) reply to Renault, Bergeaud, and Bosquet (RBB) and show that the stopping-time selection artifact RBB identify — which can mechanically create a positive event‑study path when treatment is dated at first detected use — is real but quantitatively bounded and cannot account for the observed positive association between LLM-detection and author productivity. Across multiple alternative designs in which the specific timing artifact cannot mechanically generate the result, a positive association persists; analogous pre‑ChatGPT placebo exercises return null effects.

Key Points

  • RBB’s stopping-time mechanism: dating adoption to the first month a detector flags a paper makes the month immediately before that first flag unusually low (by a factor 1−p), so post‑detection months look elevated even under a null.
  • Two problems with RBB’s original placebo approach:
  • A placebo that reproduces the event‑study shape does not prove a null — when a genuine effect exists, virtually any output‑correlated flag can partially trace the true timing.
  • In the data, random/placebo flags correlate with high‑output authors (over‑selection of adopters among placebo‑treated), so RBB’s placebos are contaminated and sit above the pure mechanical artifact (making them conservative benchmarks).
  • Calibrated rate‑matched placebo: when the random placebo is set to the detector’s realized flag rate, the detector’s event‑study still sits significantly above the placebo benchmark. Reported excesses:
    • α > 0.1 threshold (p* = 0.147): average post‑treatment excess ≈ 0.119 log points (95% CI [0.068, 0.169]).
    • α > 0.5 threshold (p* = 0.054): excess ≈ 0.183 log points (95% CI [0.121, 0.244]).
  • Four alternative designs where the RBB timing artifact cannot drive results — all continue to show a positive association while corresponding pre‑ChatGPT placebo tests are null:
  • Before‑and‑after across separated periods (date adoption in 2023, measure output 2024 vs 2022): increases ≈ +17.3% (α>0.1) and +25.1% (α>0.5).
  • Conservative control group (never‑treated + not‑yet‑treated pooled): pooled DiD ≈ 0.068 log points (95% CI [0.053, 0.083]); design attenuates true effects, so positive estimate is a lower bound.
  • Intensity (lagged) specification with no discrete adoption date: productivity regressed on lagged AI intensity yields positive, significant coefficients (e.g., average detector score lag: +0.103, p<0.001); pre‑ChatGPT placebo estimates are null/negative and insignificant.
  • Rank‑based measurement holding flag rate fixed: at fixed p, flagging top‑α papers sits above a random baseline and bottom‑α below it (e.g., 15.3% difference at p=0.2%), inconsistent with a pure p‑driven mechanical null.
  • External corroboration: an independent analysis with different data/identification finds the same sign.
  • Authors remain explicit that these associations should not be read as definitive causal effect magnitudes; dating adoption from first detection has known limitations.

Data & Methods

  • Data: arXiv author–paper time series covering the period around ChatGPT emergence; LLM detector assigns paper‑level scores and flags at thresholds (α > 0.1 and α > 0.5). Pre‑ChatGPT (2020–2022) data used for placebo tests.
  • Baseline event‑study: adopter vs non‑adopter comparison where treatment month is the author’s first flagged month.
  • Placebo strategies:
    • RBB’s random flags (shown to be contaminated by output correlations).
    • Rate‑matched random placebo: random Bernoulli flag calibrated to the detector’s realized flag rate (conservative benchmark for the timing artifact).
  • Robust alternative designs (each intended to remove or difference out the depressed reference month):
  • Before/after across separated calendar windows (adoption dated in one year, output measured in another).
  • Conservative DiD control group that pools never‑treated and not‑yet‑treated (matches the selection asymmetry so mechanical jump differences out).
  • Intensity specification: regress current productivity on lagged AI intensity (average detector score; shares of papers above thresholds), with author and quarter fixed effects.
  • Rank‑based flagging: for fixed p, compare top‑p% (highest α), bottom‑p% (lowest α), and random Bernoulli(p).
  • Placebo checks: identical procedures applied to pre‑ChatGPT data return null/insignificant results throughout.
  • Quantitative inference: event‑study coefficients, aggregated post‑treatment excesses, confidence intervals reported for key comparisons.

Implications for AI Economics

  • Empirical implication: there is a robust, positive association between detected LLM assistance and short‑term scientific productivity in this arXiv sample; the association survives multiple designs that rule out the specific stopping‑time artifact identified by RBB.
  • Methodological lessons:
    • Dating treatment from the outcome (first detected flagged paper) can induce a mechanical artifact; researchers must assess and calibrate for this (e.g., match flag rates, use lagged/intensity measures, or adopt conservative control groups).
    • Placebo designs must be checked for contamination: output‑correlated randomness can overstate the artifact and produce misleading comparisons.
    • Use multiple complementary designs (rate‑matched placebo, lagged intensity, conservative DiD, rank‑based tests, separated-period comparisons) and pre‑treatment/placebo periods to triangulate inference.
  • Policy and evaluation:
    • Early evidence suggests LLM assistance correlates with higher researcher output; policymakers and institutions evaluating AI’s research impacts should consider both statistical associations and identification robustness before drawing causal or prescriptive conclusions.
    • Because the authors refrain from claiming a causal estimate, policymakers should treat these results as indicative and prioritize further causal work (e.g., randomized encouragement, instrumenting AI use, richer heterogeneity and mechanism exploration).
  • Research agenda:
    • Investigate mechanisms (time savings vs. idea generation vs. lower quality threshold), heterogeneity across fields and career stages, and longer‑run effects (quality, citation impacts, downstream innovation).
    • Replicate across other publication venues and detector designs; develop standards for flag‑rate calibration and for measuring adoption when treatment is only observable via outcome signals.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Multiple complementary observational designs and placebo exercises consistently show a positive association and rule out the specific stopping-time artifact identified by RBB, but all analyses are non-experimental, rely on a detector with measurement error, and cannot fully exclude unobserved confounding or contemporaneous secular trends. Methods Rigorhigh — The authors directly address the critic's mechanism, calibrate placebos to the detector's realized flag rate, use conservative control groups that attenuate (not inflate) estimates, run several alternative specifications (including lagged intensity and rank-based tests) and replicate placebo checks on pre-ChatGPT data—demonstrating careful, multi-pronged robustness work. SampleAuthors and monthly paper-output on arXiv (preprints), with an LLM-detector applied to abstracts producing per-paper scores and binary flags at thresholds α>0.1 and α>0.5; empirical period mainly 2022–2024 (post-ChatGPT), with pre-ChatGPT placebo period 2020–2022; analyses use author-level monthly or quarterly productivity (paper counts) and subsets by baseline publication deciles and keywords. Themesproductivity adoption human_ai_collab IdentificationObservational adopter/non-adopter comparisons using stacked event-study / difference-in-differences; calibrated random-placebo benchmarks (matching detector flag rate and first-detection spike); conservative control group (never-treated + not-yet-treated) DiD; before-after comparison across separated years; intensity regressions of lagged AI-use share/score; rank-based comparisons holding flag rate fixed; pre-ChatGPT placebo tests. GeneralizabilitySample limited to arXiv preprint authors and fields represented there (may not generalize to industry R&D or non-preprint academic publishing), Outcome measures focus on quantity (paper counts) rather than quality, downstream impact, or economic outcomes like wages or firm productivity, LLM detector is imperfect; classification error and threshold choice can affect results, Observational design vulnerable to unobserved time-varying confounders and contemporaneous shocks (even if many checks reduce this concern), Short-to-medium run window (2022–24) — long-run effects not assessed

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Recalibrating a random placebo to the detector's realized flag rate leaves the empirical LLM-productivity association above the stopping-time benchmark. Research Productivity positive Change in author productivity relative to the first detected adoption month
Reading fidelity high
Study strength high
Average post-treatment excess of 0.119 log points at α > 0.1 and 0.183 log points at α > 0.5
0.8
The average post-treatment excess of the empirical detector over the matched random placebo was 0.119 log points at the α > 0.1 threshold. Research Productivity positive Post-treatment author productivity excess relative to the random placebo
Reading fidelity high
Study strength high
0.119 log points (95% CI [0.068, 0.169])
0.8
The average post-treatment excess of the empirical detector over the matched random placebo was 0.183 log points at the stricter α > 0.5 threshold. Research Productivity positive Post-treatment author productivity excess relative to the random placebo
Reading fidelity high
Study strength high
0.183 log points (95% CI [0.121, 0.244])
0.8
Authors classified as empirical LLM adopters were over-represented among placebo-treated authors by approximately 1.3 to 1.5 times in aggregate, even after conditioning on baseline output subgroups. Automation Exposure positive Association between empirical adopter status and placebo-treatment assignment
Reading fidelity high
Study strength medium
roughly 1.3 to 1.5 times
0.48
A before-and-after comparison found that authors classified as adopters based on 2023 detector flags increased average monthly output by 17.3% at α > 0.1 and 25.1% at α > 0.5 when comparing January–June 2024 with January–June 2022. Research Productivity positive Average monthly scientific output per author
Reading fidelity high
Study strength medium
17.3% increase at α > 0.1; 25.1% increase at α > 0.5
0.48
The corresponding before-and-after pre-ChatGPT placebo comparison returned approximately no positive effect: −0.9% at α > 0.1 and 0.0% at α > 0.5. Research Productivity null_result Change in average monthly scientific output in the pre-ChatGPT placebo period
Reading fidelity high
Study strength medium
−0.9% at α > 0.1; 0.0% at α > 0.5
0.48
A difference-in-differences specification using never-treated and not-yet-treated authors as controls produced a pooled positive association of 0.068 log points. Research Productivity positive Change in author productivity in a 2×2 difference-in-differences design
Reading fidelity high
Study strength high
0.068 log points (95% CI [0.053, 0.083])
0.8
Prior-period AI-use intensity was positively and significantly associated with current productivity for all three intensity measures in the empirical data. Research Productivity positive Current quarterly author productivity
Reading fidelity high
Study strength high
Average detector score: +0.103; share α > 0.1: +0.031; share α > 0.5: +0.071
0.8
The corresponding pre-ChatGPT placebo regressions showed negative and statistically insignificant associations for all three prior-period AI-use intensity measures. Research Productivity null_result Current quarterly author productivity in pre-ChatGPT placebo data
Reading fidelity high
Study strength medium
−0.086, −0.023, and −0.101; p = 0.14, 0.17, and 0.15
0.48
At a fixed flagging rate of p = 0.2%, the top-α and bottom-α flagging rules differed in their estimated productivity changes by 15.3%, with the top-α rule above and the bottom-α rule below the random baseline. Research Productivity mixed Change in author productivity under alternative paper-flagging rules
Reading fidelity high
Study strength medium
15.3% difference at flagging rate p = 0.2%
0.48
Under the authors' simulated null, the stopping-time artifact produced a single-month productivity spike at the adoption month but no sustained productivity shift. Research Productivity null_result Simulated author productivity before and after mechanically dated adoption
Reading fidelity high
Study strength medium
approximately zero under the simulated null
0.48

Notes