0 cumulative citations
View corpus contextDetected LLM use on arXiv associates with a sustained rise in authors' paper output that exceeds what a stopping-time timing artifact would produce; calibrated placebos and several conservative designs preserve the positive association though causality is not claimed.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.
Summary
Main Finding
Kusumegi et al. (2026) reply to Renault, Bergeaud, and Bosquet (RBB) and show that the stopping-time selection artifact RBB identify — which can mechanically create a positive event‑study path when treatment is dated at first detected use — is real but quantitatively bounded and cannot account for the observed positive association between LLM-detection and author productivity. Across multiple alternative designs in which the specific timing artifact cannot mechanically generate the result, a positive association persists; analogous pre‑ChatGPT placebo exercises return null effects.
Key Points
- RBB’s stopping-time mechanism: dating adoption to the first month a detector flags a paper makes the month immediately before that first flag unusually low (by a factor 1−p), so post‑detection months look elevated even under a null.
- Two problems with RBB’s original placebo approach:
- A placebo that reproduces the event‑study shape does not prove a null — when a genuine effect exists, virtually any output‑correlated flag can partially trace the true timing.
- In the data, random/placebo flags correlate with high‑output authors (over‑selection of adopters among placebo‑treated), so RBB’s placebos are contaminated and sit above the pure mechanical artifact (making them conservative benchmarks).
- Calibrated rate‑matched placebo: when the random placebo is set to the detector’s realized flag rate, the detector’s event‑study still sits significantly above the placebo benchmark. Reported excesses:
- α > 0.1 threshold (p* = 0.147): average post‑treatment excess ≈ 0.119 log points (95% CI [0.068, 0.169]).
- α > 0.5 threshold (p* = 0.054): excess ≈ 0.183 log points (95% CI [0.121, 0.244]).
- Four alternative designs where the RBB timing artifact cannot drive results — all continue to show a positive association while corresponding pre‑ChatGPT placebo tests are null:
- Before‑and‑after across separated periods (date adoption in 2023, measure output 2024 vs 2022): increases ≈ +17.3% (α>0.1) and +25.1% (α>0.5).
- Conservative control group (never‑treated + not‑yet‑treated pooled): pooled DiD ≈ 0.068 log points (95% CI [0.053, 0.083]); design attenuates true effects, so positive estimate is a lower bound.
- Intensity (lagged) specification with no discrete adoption date: productivity regressed on lagged AI intensity yields positive, significant coefficients (e.g., average detector score lag: +0.103, p<0.001); pre‑ChatGPT placebo estimates are null/negative and insignificant.
- Rank‑based measurement holding flag rate fixed: at fixed p, flagging top‑α papers sits above a random baseline and bottom‑α below it (e.g., 15.3% difference at p=0.2%), inconsistent with a pure p‑driven mechanical null.
- External corroboration: an independent analysis with different data/identification finds the same sign.
- Authors remain explicit that these associations should not be read as definitive causal effect magnitudes; dating adoption from first detection has known limitations.
Data & Methods
- Data: arXiv author–paper time series covering the period around ChatGPT emergence; LLM detector assigns paper‑level scores and flags at thresholds (α > 0.1 and α > 0.5). Pre‑ChatGPT (2020–2022) data used for placebo tests.
- Baseline event‑study: adopter vs non‑adopter comparison where treatment month is the author’s first flagged month.
- Placebo strategies:
- RBB’s random flags (shown to be contaminated by output correlations).
- Rate‑matched random placebo: random Bernoulli flag calibrated to the detector’s realized flag rate (conservative benchmark for the timing artifact).
- Robust alternative designs (each intended to remove or difference out the depressed reference month):
- Before/after across separated calendar windows (adoption dated in one year, output measured in another).
- Conservative DiD control group that pools never‑treated and not‑yet‑treated (matches the selection asymmetry so mechanical jump differences out).
- Intensity specification: regress current productivity on lagged AI intensity (average detector score; shares of papers above thresholds), with author and quarter fixed effects.
- Rank‑based flagging: for fixed p, compare top‑p% (highest α), bottom‑p% (lowest α), and random Bernoulli(p).
- Placebo checks: identical procedures applied to pre‑ChatGPT data return null/insignificant results throughout.
- Quantitative inference: event‑study coefficients, aggregated post‑treatment excesses, confidence intervals reported for key comparisons.
Implications for AI Economics
- Empirical implication: there is a robust, positive association between detected LLM assistance and short‑term scientific productivity in this arXiv sample; the association survives multiple designs that rule out the specific stopping‑time artifact identified by RBB.
- Methodological lessons:
- Dating treatment from the outcome (first detected flagged paper) can induce a mechanical artifact; researchers must assess and calibrate for this (e.g., match flag rates, use lagged/intensity measures, or adopt conservative control groups).
- Placebo designs must be checked for contamination: output‑correlated randomness can overstate the artifact and produce misleading comparisons.
- Use multiple complementary designs (rate‑matched placebo, lagged intensity, conservative DiD, rank‑based tests, separated-period comparisons) and pre‑treatment/placebo periods to triangulate inference.
- Policy and evaluation:
- Early evidence suggests LLM assistance correlates with higher researcher output; policymakers and institutions evaluating AI’s research impacts should consider both statistical associations and identification robustness before drawing causal or prescriptive conclusions.
- Because the authors refrain from claiming a causal estimate, policymakers should treat these results as indicative and prioritize further causal work (e.g., randomized encouragement, instrumenting AI use, richer heterogeneity and mechanism exploration).
- Research agenda:
- Investigate mechanisms (time savings vs. idea generation vs. lower quality threshold), heterogeneity across fields and career stages, and longer‑run effects (quality, citation impacts, downstream innovation).
- Replicate across other publication venues and detector designs; develop standards for flag‑rate calibration and for measuring adoption when treatment is only observable via outcome signals.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Recalibrating a random placebo to the detector's realized flag rate leaves the empirical LLM-productivity association above the stopping-time benchmark. Research Productivity | positive | Change in author productivity relative to the first detected adoption month |
Reading fidelity
high
Study strength
high
|
Average post-treatment excess of 0.119 log points at α > 0.1 and 0.183 log points at α > 0.5
|
| The average post-treatment excess of the empirical detector over the matched random placebo was 0.119 log points at the α > 0.1 threshold. Research Productivity | positive | Post-treatment author productivity excess relative to the random placebo |
Reading fidelity
high
Study strength
high
|
0.119 log points (95% CI [0.068, 0.169])
|
| The average post-treatment excess of the empirical detector over the matched random placebo was 0.183 log points at the stricter α > 0.5 threshold. Research Productivity | positive | Post-treatment author productivity excess relative to the random placebo |
Reading fidelity
high
Study strength
high
|
0.183 log points (95% CI [0.121, 0.244])
|
| Authors classified as empirical LLM adopters were over-represented among placebo-treated authors by approximately 1.3 to 1.5 times in aggregate, even after conditioning on baseline output subgroups. Automation Exposure | positive | Association between empirical adopter status and placebo-treatment assignment |
Reading fidelity
high
Study strength
medium
|
roughly 1.3 to 1.5 times
|
| A before-and-after comparison found that authors classified as adopters based on 2023 detector flags increased average monthly output by 17.3% at α > 0.1 and 25.1% at α > 0.5 when comparing January–June 2024 with January–June 2022. Research Productivity | positive | Average monthly scientific output per author |
Reading fidelity
high
Study strength
medium
|
17.3% increase at α > 0.1; 25.1% increase at α > 0.5
|
| The corresponding before-and-after pre-ChatGPT placebo comparison returned approximately no positive effect: −0.9% at α > 0.1 and 0.0% at α > 0.5. Research Productivity | null_result | Change in average monthly scientific output in the pre-ChatGPT placebo period |
Reading fidelity
high
Study strength
medium
|
−0.9% at α > 0.1; 0.0% at α > 0.5
|
| A difference-in-differences specification using never-treated and not-yet-treated authors as controls produced a pooled positive association of 0.068 log points. Research Productivity | positive | Change in author productivity in a 2×2 difference-in-differences design |
Reading fidelity
high
Study strength
high
|
0.068 log points (95% CI [0.053, 0.083])
|
| Prior-period AI-use intensity was positively and significantly associated with current productivity for all three intensity measures in the empirical data. Research Productivity | positive | Current quarterly author productivity |
Reading fidelity
high
Study strength
high
|
Average detector score: +0.103; share α > 0.1: +0.031; share α > 0.5: +0.071
|
| The corresponding pre-ChatGPT placebo regressions showed negative and statistically insignificant associations for all three prior-period AI-use intensity measures. Research Productivity | null_result | Current quarterly author productivity in pre-ChatGPT placebo data |
Reading fidelity
high
Study strength
medium
|
−0.086, −0.023, and −0.101; p = 0.14, 0.17, and 0.15
|
| At a fixed flagging rate of p = 0.2%, the top-α and bottom-α flagging rules differed in their estimated productivity changes by 15.3%, with the top-α rule above and the bottom-α rule below the random baseline. Research Productivity | mixed | Change in author productivity under alternative paper-flagging rules |
Reading fidelity
high
Study strength
medium
|
15.3% difference at flagging rate p = 0.2%
|
| Under the authors' simulated null, the stopping-time artifact produced a single-month productivity spike at the adoption month but no sustained productivity shift. Research Productivity | null_result | Simulated author productivity before and after mechanically dated adoption |
Reading fidelity
high
Study strength
medium
|
approximately zero under the simulated null
|