0 cumulative citations
View corpus contextAI will first enlarge and then shrink research teams: a formal model finds teams peak when automated task coverage nears a high threshold (median ≈91%), forecasting early peaks in the most codifiable fields (algebra and theoretical computer science) around 2027–2029; current data through 2025 show no detectable change, consistent with limited effective adoption.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
1 cumulative citations
View corpus contextArtificial intelligence is associated with larger research teams, yet in mathematics, among the most codifiable fields, individual researchers working with AI now produce research-grade results. A span-of-control model reconciles these observations. AI lowers execution cost, which expands laboratory scale, and automates codifiable tasks, which lowers the member share of each unit. Team size is therefore quasi-concave in AI capability, with at most one peak. The model predicts that a fully codifiable team peaks when effective automation coverage reaches a closed-form threshold, typically near complete coverage, and, among fields with shared primitives that possess an interior peak, those with less irreducibly human task content peak first. Under explicit priors, the 90 percent forecast intervals for the fully codifiable peak span 2026 to 2030.
Summary
Main Finding
AI has two opposing effects on research teams: it (i) lowers execution costs and so expands laboratory scale, and (ii) automates codifiable execution tasks and so reduces the member share per unit of execution. Modeling those forces yields a single-peaked (quasi-concave) relationship between AI capability and research-team size: teams can grow and then shrink as AI capability rises, with at most one turning point. For fully codifiable fields the model gives a closed-form threshold for automated-task coverage at the peak (median s* ≈ 0.91 under the paper’s priors). Mapping current capability trends into this framework produces a falsifiable dated forecast: the fully codifiable peak is most likely around 2027 (median 2027.2 under instantaneous adoption), or about 2028.5 after plausible diffusion lags; if long-run adoption stalls below high levels, a peak may never occur.
Key Points
- Mechanism: leader judgment is fixed (one principal investigator); execution is assembled from tasks. A share ϕ of tasks is irreducibly human; the rest is codifiable and progressively automatable as capability a rises.
- Team size decomposition: team members m(a) = X(a) · [1 − σ(a)], where X (laboratory scale) rises when execution costs fall, and σ(a) is the automated share of codifiable tasks. Total team N = 1 (leader) + m*.
- Shape result: m(a) and N(a) are quasi‑concave in capability a — at most one rise and one fall.
- Closed-form peak condition (fully codifiable case, ϕ=0):
- Automated-share-at-peak: s* = (β − r) / [β(1 − r)], where β is the execution elasticity and r = pA/q is the AI task price relative to human task price.
- Field threshold for interior peak: ϕ = r(1 − β) / [β(1 − r)] and s + ϕ* = 1.
- Intuition: the peak occurs when increased scale no longer offsets substitution away from human members.
- Empirical/short-run evidence (through 2025): no detectable change in author-counts attributable to AI exposure. Survey evidence dates routine use of coding agents among quantitative social scientists to late December 2025 (≈20% regular users by early 2026), implying little effective adoption through 2025 — consistent with the model’s prediction of no detectable change pre-diffusion.
- Watchlist (fields most likely to peak first): the most codifiable fields (lowest estimated ϕ) — notably algebra & number theory and theoretical computer science (zero-score proxies), plus a set of seven lowest-proxy fields in the paper’s benchmark — should be monitored for first turning points in author counts.
- Forecast sensitivity:
- Under priors used, s* median ≈ 0.91 (90% interval ≈ [0.75, 0.97]).
- Capability-clock Monte Carlo (METR bridge) median peak year 2027.2 (90% interval [2026.4, 2028.1]); with 0.5–2.0 year diffusion lag median 2028.5 ([2027.4, 2029.6]).
- If long-run adoption ceiling drawn U(0.8,1.0), a peak occurs in ≈50% of draws (conditional median ~2028.6).
- Important caveats: results depend on (i) mapping software capability to task automation (G); (ii) adoption rates; (iii) fixed task prices in the comparative static; and (iv) the proxy used to measure field irreducible human-task share ϕ.
Data & Methods
Model - Setup: Each project output yi = Ei^β (C ti)^(1−β) (Ei = execution, ti = leader attention, C = leader judgment). Optimal attention allocation yields span-of-control result: ti ∝ Ei and aggregate output ∝ C^(1−β) X^β where X = ΣEi. - Tasks and automation: - Share ϕ ∈ [0,1) is irreducibly human. - Codifiable share 1 − ϕ is automatable up to σ(a) = (1 − ϕ) G(a), with G(a) increasing to 1 as capability a → ∞. - Unit execution cost c(a) = q − (q − pA) σ(a), where q is human task cost and pA < q is AI task cost (so r = pA/q ∈ (0,1)). - Equilibrium scale: X(a) = C [β / c(a)]^(1/(1−β)); member employment m(a) = X(a) · [1 − σ(a)]; N(a) = 1 + m(a). - Main analytical result: d ln m/ds has numerator linear-decreasing in s, so sign changes at most once ⇒ quasi‑concavity. The closed-form s and ϕ above derive from setting that numerator to zero.
Empirical evidence through 2025 - Field-level proxy for ϕ: manually scored 120 pre-period abstracts per OpenAlex subfield (30 subfields) from 2015 & 2019, scoring per-article execution mode (0 theory/simulation/software/secondary-data; 1 bench/field/clinical/fabrication; 0.5 mixed). Field means bϕf used as ordinal proxies for model’s ϕ. - Panel of PIs: 2,279 principal investigators (2015–2019 IDs) tracked monthly across 13 preprint servers (36,186 preprints). Outcomes: log author count per paper; auxiliary margins: annual papers, credited authorships per investigator. Regressions include investigator × month, field-by-month, and server fixed effects; interaction between predetermined exposure and mapped capability tested. - Survey: Feb–Mar 2026, 1,260 quantitative social scientists; ~20% reported regular use of coding agents; surge dated to late Dec 2025. - Results: no statistically significant author-count response through Dec 2025; within-field computational papers gain modestly after 2022 but overall null; consistent with pre-diffusion baseline.
Forecast Monte Carlo - Capability bridge: uses METR (software doubling / capability-to-software horizon) trend (Kwa et al., 2025) and functional form G(a) = a / (a + h50) where h50 is median task duration. - Priors propagated (20,000 draws): β ∼ Uniform(0.4,0.8); r ∼ Uniform(0.05,0.20); h50 log-uniform 2–24 hours; bootstrap of post-2023 METR trend for doubling-time uncertainty. Scenario A: instantaneous adoption; Scenario B: add uniform lag ℓ ∈ [0.5, 2.0] years for diffusion/output delays. Scenario C: draw long-run adoption ceiling U(0.8,1.0) (peak conditional on ceiling). - Output: distribution of predicted peak years and peak probabilities under ceilings.
Implications for AI Economics
- Nonmonotone labor effects: automation can first increase employment (via scale) then reduce it (via substitution). This creates transition dynamics with potentially large distributional consequences for researchers’ careers and earnings depending on transition speed.
- Field heterogeneity matters: fields with lower irreducible human-task shares (more codifiable) should be earliest affected — they are the “watchlist” where peaks should appear first. Policy and institutional responses should be targeted by field (training, hiring/funding adjustments).
- Monitoring priorities for forecasting and policy:
- Effective automated coverage (capability × adoption) relative to s*.
- AI task price r (pA/q): cheaper AI shifts thresholds and can eliminate the interior peak.
- Adoption rates and diffusion lags (surveys, usage telemetry).
- Field task-mix measurement (better estimates of ϕ than ordinal proxies).
- Research labor-market implications:
- If model assumptions hold, long-run equilibrium for fully codifiable fields is one leader overseeing machine execution — implications for training pipelines and human skill formation (if execution is where judgment is learned, strong automation could erode the formation of future leaders).
- Transitional welfare effects depend on speed of change and frictions; faster transitions could impose larger lifetime-earnings costs on displaced researchers.
- Testing and falsification: The framework yields concrete, falsifiable tests — notably whether author counts in the lowest-ϕ fields (algebra & number theory; theoretical CS) peak by the forecasted window given confirmed crossing of s*. Observed peaks earlier than predicted would reject the model plus its bridge assumptions; absence of peaks after confirmed crossing would also reject the joint hypothesis.
- Model limitations to bear in mind:
- Results use fixed relative prices and a Cobb–Douglas execution technology; endogenous price declines or different substitution elasticities can change dynamics (multiple turning points possible).
- The mapping from software capability to task automation (G) and the proxy mapping from article scores to ϕ are key empirical inputs and sources of uncertainty.
- Funding-driven headcounts or production choices (more projects at constant authors-per-paper) can mask production-side employment changes.
If you want, I can: - Extract the specific closed-form formulas and short intuition for s and ϕ into a one-page cheat sheet. - Produce a concise monitoring dashboard listing the exact indicators (with suggested data sources) to watch to test the paper’s predictions over 2026–2030.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In the model, research-team size is quasi-concave in AI capability: it can rise at most once and fall at most once, with a single turning point from increasing to decreasing. Team Performance | mixed | Research-team size as a function of AI capability |
Reading fidelity
high
Study strength
high
|
not reported
|
| A fully codifiable research team reaches its peak size when its effective automated task share reaches the threshold s* = (β − r)/(β(1 − r)), where r is the AI-to-member task-cost ratio. Team Performance | mixed | Peak research-team size as a function of automated task coverage |
Reading fidelity
high
Study strength
high
|
s* = (β − r)/(β(1 − r))
|
| Under the paper's prior distribution for model parameters, the peak automated-task share for a fully codifiable team has a median of 0.91 and a 90% interval from 0.75 to 0.97. Automation Exposure | other | Automated task coverage at the research-team-size peak |
Reading fidelity
high
Study strength
medium
|
n=20000
median 0.91; 90% interval [0.75, 0.97]
|
| Among fields sharing the model's other primitives and having an interior peak, fields with smaller irreducible human-task shares are predicted to peak earlier. Task Allocation | negative | Timing of the research-team-size peak by field |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The field-level empirical estimates do not detect a statistically significant relationship between physical exposure and post-2022 author-count growth after including field-specific trends. Team Performance | null_result | Author count per research paper |
Reading fidelity
high
Study strength
low
|
n=3600
none statistically distinguishable from zero
|
| At the researcher level, the interaction between predetermined physical exposure and mapped AI capability is 0.19, with a standard error of 0.26, and is statistically indistinguishable from zero through December 2025. Team Performance | null_result | Log author count per preprint |
Reading fidelity
high
Study strength
medium
|
n=2279
interaction = 0.19; standard error = 0.26
|
| The same researcher-level null result appears when laboratory scale is measured using annual papers and total credited authorships per investigator. Research Productivity | null_result | Annual papers and total credited authorships per investigator |
Reading fidelity
high
Study strength
medium
|
n=2279
|
| In a February–March 2026 survey of quantitative social scientists, 20% reported regularly using coding agents, with the surge in use dating to late December 2025. Adoption Rate | positive | Regular adoption of coding agents |
Reading fidelity
high
Study strength
low
|
n=1260
20 percent used coding agents regularly
|
| Under the paper's capability-to-task benchmark and priors, a fully codifiable research team is forecast to peak at a median date of 2027.2 on the capability clock, with a 90% interval of 2026.4 to 2028.1. Team Performance | negative | Calendar date of the fully codifiable research-team-size peak |
Reading fidelity
high
Study strength
medium
|
n=20000
median 2027.2; 90 percent interval [2026.4, 2028.1]
|
| After adding an adoption-and-output lag of 0.5 to 2.0 years, the forecasted median peak date shifts to 2028.5, with a 90% interval from 2027.4 to 2029.6. Team Performance | negative | Calendar date of the fully codifiable research-team-size peak after adoption lag |
Reading fidelity
high
Study strength
low
|
n=20000
median 2028.5; 90 percent interval [2027.4, 2029.6]
|
| If long-run adoption reaches only 80% to 100% of codifiable tasks, the model produces a peak in approximately half of simulation draws; below that range, a peak does not occur. Adoption Rate | mixed | Occurrence of a research-team-size peak under an adoption ceiling |
Reading fidelity
high
Study strength
medium
|
n=20000
peak occurs in 50 percent of draws
|