0 cumulative citations
View corpus contextA reanalysis of METR's capability series finds no strong evidence for continued exponential growth and shows a plausible inflection point has already passed; a two-component model of base and reasoning skills further implies near-term slowing under reasonable assumptions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Rapidly increasing AI capabilities have substantial real-world consequences, ranging from AI safety concerns to labor market consequences. The Model Evaluation & Threat Research (METR) report argues that AI capabilities have exhibited exponential growth since 2019. In this note, we argue that the data does not support exponential growth, even in shorter-term horizons. Whereas the METR study claims that fitting sigmoid/logistic curves results in inflection points far in the future, we fit a sigmoid curve to their current data and find that the inflection point has already passed. In addition, we propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, exhibiting individual rates of improvement. We prove that this model supports our hypothesis that AI capabilities will exhibit an inflection point in the near future. Our goal is not to establish a rigorous forecast of our own, but to highlight the fragility of existing forecasts of exponential growth.
Summary
Main Finding
The authors challenge METR's claim of sustained exponential growth in frontier AI capabilities. They show that (a) a simple sigmoid fit to METR’s 50% model-horizon data is consistent with an inflection (slowing) already or imminently occurring, and (b) a plausible multiplicative model that decomposes overall capability into base LLM capability and reasoning capability (each following sigmoid-like growth) naturally produces an early exponential-looking phase followed by plateauing once component inflection points are crossed. In-sample fits favor the sigmoid-linked multiplicative model over pure exponential links, implying that continued exponential growth is not the only (and may not be the most supported) interpretation of current data.
Key Points
- METR background: METR (Kwa et al., 2025) introduced the 50% model-horizon metric (unbounded) and reported that frontier capabilities double roughly every seven months since 2019, based on regressing log(horizon) on release date (i.e., an exponential trend, R^2 ≈ 0.98).
- Reanalysis: Fitting a simple sigmoid to METR’s reported 50% horizon points yields an inflection point consistent with slowing growth (authors cite ~2025-06-06 from a simple fit).
- Multiplicative hypothesis: The paper models overall capability as multiplicative: h_model = γ1 · h_base(d) · (1 + γ2 · h_reasoning(d) · 1{kthinking}), separating base pretraining progress from post-training reasoning improvements.
- Link functions tested: sigmoid, exponential, and B-spline for the time dependence of base and reasoning components.
- Theoretical result: Product of sigmoids with staggered inflection times exhibits (i) exponential growth before the first inflection, (ii) an amplified/compound growth region between inflection points, and (iii) plateau after the last inflection — explaining why staggered innovations can look exponential for some time.
- Empirical fit: Using METR’s raw task data (HCAST, RE-Bench, SWAA; 15 SOTA models), estimated via probabilistic likelihood (Stan) with weak priors:
- Sigmoid-link multiplicative model had lower in-sample MSE than exponential-link multiplicative model and METR’s exponential fit (table of MSEs reported).
- Estimated inflection dates in their fitted multiplicative model: base inflection ≈ 2024-11-21; reasoning inflection projected ≈ mid-2026 (authors present a near-term inflection for reasoning).
- Caveats: All results are in-sample (small number of frontier models), different loss functions between models, and forecasts are fragile — data alone cannot definitively rule out continued exponential growth.
Data & Methods
- Data: METR public dataset (170 tasks across HCAST, RE-Bench, SWAA; 28 evaluated models, focusing on 15 SOTA/frontier models and their release dates).
- Metric: 50% model-horizon time h_model (difficulty/time threshold at which a model achieves 50% success probability). METR estimates h_model by fitting p_model = σ((log h_model − log t_task) · β_model).
- Baseline (METR) approach: Linear regression on log h_model versus model release date d_model (implies exponential time trend).
- Authors’ approach:
- Structural model: h_model = γ1 · h_base(d) · (1 + γ2 · h_reasoning(d)), with h_reasoning active only when reasoning post-training is used.
- Link forms: tested sigmoid (σ(α d + β)), exponential (exp(α d+β)), and B-spline (degree-5, two breakpoints) for b(d) and r(d).
- Estimation: Maximum (log-)likelihood / Bayesian-style estimation with Stan, weakly informative N(0,10^2) priors for many parameters, positivity constraints for scale parameters; spline coefficients regularized with random-walk priors to avoid overfitting.
- Evaluation: In-sample MSE on predicted versus observed h_model and inspection of inferred inflection points; visual comparison of projections out to 2029.
- Reported fit summary (in-sample MSE on h_model): Sigmoid-link multiplicative model ≈ 203.7; METR exponential curve ≈ 339.9; exponential-link multiplicative ≈ 2874.7; a separately fitted simple sigmoid curve to METR points reported much lower MSE (~27.4) for that particular fit (different loss target).
Implications for AI Economics
- Forecast uncertainty matters: Economic and policy decisions driven by expectations of rapid, continued exponential capability growth (e.g., labor market adjustment, education choices, regulation) are sensitive to model assumptions. The paper shows those assumptions are fragile and that plateauing is a plausible alternative.
- Timing of displacement and complementary investments: If capabilities plateau soon (or growth slows), large-scale, rapid displacement of skilled labor may be delayed or diminished — affecting optimal timing for retraining, social insurance, and capital reallocation decisions.
- Importance of component decomposition: Treating capabilities as multiplicative components (base model × reasoning, etc.) suggests that economic models of automation should consider heterogeneous, interacting technologies with staggered innovation timing — not a single homogeneous exponential trend.
- Value of robustness and scenario analysis: Policymakers and firms should plan for a range of trajectories (continued exponential, transient exponential then plateau, episodic breakthroughs) and stress-test investments and regulations against these scenarios.
- Role of breakthroughs: Under the multiplicative view, sustaining exponential growth depends on further major innovations (new components); economic forecasts should incorporate probabilities and expected impacts of such breakthroughs rather than mechanically extrapolating past growth.
- Research and data needs: Better, larger-scale, and temporal datasets on frontier models (and metric choices) are crucial. Economic modeling of AI impact should incorporate model uncertainty, small-sample fragility, and diagnostics for inflection/structural changes.
Limitations to keep in mind: the paper’s results are in-sample and rely on a small frontier-model sample; loss-function and model-specification choices affect conclusions; data cannot definitively rule out exponential growth — the contribution is to highlight a compelling, alternative interpretation that matters for economic inference and policy.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The METR report argues that AI capabilities have exhibited exponential growth since 2019. Research Productivity | positive | AI capabilities growth over time |
Reading fidelity
high
Study strength
high
|
not reported
|
| The data does not support exponential growth (of AI capabilities), even in shorter-term horizons. Research Productivity | negative | AI capabilities growth over time |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The METR study claims that fitting sigmoid/logistic curves results in inflection points far in the future. Research Productivity | positive | timing of inflection point in fitted capability trajectories |
Reading fidelity
high
Study strength
high
|
not reported
|
| When we fit a sigmoid curve to METR's current data, the inflection point has already passed. Research Productivity | negative | timing of inflection point in fitted capability trajectories |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, each exhibiting individual rates of improvement. Research Productivity | mixed | structure of capability growth (base vs reasoning components and their rates) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We prove that this decomposed model supports the hypothesis that AI capabilities will exhibit an inflection point in the near future. Research Productivity | negative | presence and timing of an inflection point in AI capability trajectory implied by the model |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The goal of the note is to highlight the fragility of existing forecasts of exponential growth in AI capabilities rather than to establish a rigorous forecast of our own. Research Productivity | mixed | robustness/fragility of exponential-growth forecasts |
Reading fidelity
high
Study strength
low
|
not reported
|