The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A reanalysis of METR's capability series finds no strong evidence for continued exponential growth and shows a plausible inflection point has already passed; a two-component model of base and reasoning skills further implies near-term slowing under reasonable assumptions.

Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
Haosen Ge, Hamsa Bastani, Osbert Bastani · February 04, 2026
arxiv commentary medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Haosen Ge unresolved corpus identity
  2. Hamsa Bastani unresolved corpus identity
  3. Osbert Bastani unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Haosen Ge provider ID
  2. Hamsa Bastani provider ID
  3. O. Bastani provider ID
The METR data do not robustly support exponential capability growth; fitting sigmoids to those data suggests an inflection may already have occurred, and a two-component model (base vs reasoning) mathematically produces an imminent inflection under plausible parameter values.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Rapidly increasing AI capabilities have substantial real-world consequences, ranging from AI safety concerns to labor market consequences. The Model Evaluation & Threat Research (METR) report argues that AI capabilities have exhibited exponential growth since 2019. In this note, we argue that the data does not support exponential growth, even in shorter-term horizons. Whereas the METR study claims that fitting sigmoid/logistic curves results in inflection points far in the future, we fit a sigmoid curve to their current data and find that the inflection point has already passed. In addition, we propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, exhibiting individual rates of improvement. We prove that this model supports our hypothesis that AI capabilities will exhibit an inflection point in the near future. Our goal is not to establish a rigorous forecast of our own, but to highlight the fragility of existing forecasts of exponential growth.

Summary

Main Finding

The authors challenge METR's claim of sustained exponential growth in frontier AI capabilities. They show that (a) a simple sigmoid fit to METR’s 50% model-horizon data is consistent with an inflection (slowing) already or imminently occurring, and (b) a plausible multiplicative model that decomposes overall capability into base LLM capability and reasoning capability (each following sigmoid-like growth) naturally produces an early exponential-looking phase followed by plateauing once component inflection points are crossed. In-sample fits favor the sigmoid-linked multiplicative model over pure exponential links, implying that continued exponential growth is not the only (and may not be the most supported) interpretation of current data.

Key Points

  • METR background: METR (Kwa et al., 2025) introduced the 50% model-horizon metric (unbounded) and reported that frontier capabilities double roughly every seven months since 2019, based on regressing log(horizon) on release date (i.e., an exponential trend, R^2 ≈ 0.98).
  • Reanalysis: Fitting a simple sigmoid to METR’s reported 50% horizon points yields an inflection point consistent with slowing growth (authors cite ~2025-06-06 from a simple fit).
  • Multiplicative hypothesis: The paper models overall capability as multiplicative: h_model = γ1 · h_base(d) · (1 + γ2 · h_reasoning(d) · 1{kthinking}), separating base pretraining progress from post-training reasoning improvements.
  • Link functions tested: sigmoid, exponential, and B-spline for the time dependence of base and reasoning components.
  • Theoretical result: Product of sigmoids with staggered inflection times exhibits (i) exponential growth before the first inflection, (ii) an amplified/compound growth region between inflection points, and (iii) plateau after the last inflection — explaining why staggered innovations can look exponential for some time.
  • Empirical fit: Using METR’s raw task data (HCAST, RE-Bench, SWAA; 15 SOTA models), estimated via probabilistic likelihood (Stan) with weak priors:
    • Sigmoid-link multiplicative model had lower in-sample MSE than exponential-link multiplicative model and METR’s exponential fit (table of MSEs reported).
    • Estimated inflection dates in their fitted multiplicative model: base inflection ≈ 2024-11-21; reasoning inflection projected ≈ mid-2026 (authors present a near-term inflection for reasoning).
  • Caveats: All results are in-sample (small number of frontier models), different loss functions between models, and forecasts are fragile — data alone cannot definitively rule out continued exponential growth.

Data & Methods

  • Data: METR public dataset (170 tasks across HCAST, RE-Bench, SWAA; 28 evaluated models, focusing on 15 SOTA/frontier models and their release dates).
  • Metric: 50% model-horizon time h_model (difficulty/time threshold at which a model achieves 50% success probability). METR estimates h_model by fitting p_model = σ((log h_model − log t_task) · β_model).
  • Baseline (METR) approach: Linear regression on log h_model versus model release date d_model (implies exponential time trend).
  • Authors’ approach:
    • Structural model: h_model = γ1 · h_base(d) · (1 + γ2 · h_reasoning(d)), with h_reasoning active only when reasoning post-training is used.
    • Link forms: tested sigmoid (σ(α d + β)), exponential (exp(α d+β)), and B-spline (degree-5, two breakpoints) for b(d) and r(d).
    • Estimation: Maximum (log-)likelihood / Bayesian-style estimation with Stan, weakly informative N(0,10^2) priors for many parameters, positivity constraints for scale parameters; spline coefficients regularized with random-walk priors to avoid overfitting.
    • Evaluation: In-sample MSE on predicted versus observed h_model and inspection of inferred inflection points; visual comparison of projections out to 2029.
  • Reported fit summary (in-sample MSE on h_model): Sigmoid-link multiplicative model ≈ 203.7; METR exponential curve ≈ 339.9; exponential-link multiplicative ≈ 2874.7; a separately fitted simple sigmoid curve to METR points reported much lower MSE (~27.4) for that particular fit (different loss target).

Implications for AI Economics

  • Forecast uncertainty matters: Economic and policy decisions driven by expectations of rapid, continued exponential capability growth (e.g., labor market adjustment, education choices, regulation) are sensitive to model assumptions. The paper shows those assumptions are fragile and that plateauing is a plausible alternative.
  • Timing of displacement and complementary investments: If capabilities plateau soon (or growth slows), large-scale, rapid displacement of skilled labor may be delayed or diminished — affecting optimal timing for retraining, social insurance, and capital reallocation decisions.
  • Importance of component decomposition: Treating capabilities as multiplicative components (base model × reasoning, etc.) suggests that economic models of automation should consider heterogeneous, interacting technologies with staggered innovation timing — not a single homogeneous exponential trend.
  • Value of robustness and scenario analysis: Policymakers and firms should plan for a range of trajectories (continued exponential, transient exponential then plateau, episodic breakthroughs) and stress-test investments and regulations against these scenarios.
  • Role of breakthroughs: Under the multiplicative view, sustaining exponential growth depends on further major innovations (new components); economic forecasts should incorporate probabilities and expected impacts of such breakthroughs rather than mechanically extrapolating past growth.
  • Research and data needs: Better, larger-scale, and temporal datasets on frontier models (and metric choices) are crucial. Economic modeling of AI impact should incorporate model uncertainty, small-sample fragility, and diagnostics for inflection/structural changes.

Limitations to keep in mind: the paper’s results are in-sample and rely on a small frontier-model sample; loss-function and model-specification choices affect conclusions; data cannot definitively rule out exponential growth — the contribution is to highlight a compelling, alternative interpretation that matters for economic inference and policy.

Assessment

Paper Typecommentary Evidence Strengthmedium — The note presents empirical model fits to the METR capability time series and a formal theoretical decomposition that produces an inflection under plausible parameterizations; this provides a credible counterexample to blanket exponential claims but does not provide strong, robust empirical validation (short time series, few datapoints, single data source, and sensitivity to model choice). Methods Rigormedium — The authors fit sigmoid/logistic curves to published capability data and derive a theoretically motivated two-component model with analytic properties; however, the analysis lacks extensive robustness checks, out-of-sample validation, formal model selection, and deeper sensitivity analysis of parameter choices and measurement issues in the underlying dataset. SampleUses the METR report's capability time series (benchmark performance measures summarized by METR) covering roughly 2019 to the present; fits are performed on that published dataset and the proposed model is analyzed mathematically rather than estimated on multiple independent datasets. Themesinnovation governance GeneralizabilityRelies entirely on METR's selected benchmarks which may not represent overall AI capabilities relevant to economic outcomes, Short time series (since ~2019) with relatively few datapoints limits extrapolation reliability, Model conclusions sensitive to functional-form and parameter choices; other plausible models could imply different trajectories, Benchmarks measure technical performance, not downstream economic impact (e.g., productivity, adoption, labor effects), External factors (policy, compute availability, data access, firm incentives) not modeled and could alter trajectories

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The METR report argues that AI capabilities have exhibited exponential growth since 2019. Research Productivity positive AI capabilities growth over time
Reading fidelity high
Study strength high
not reported
0.1
The data does not support exponential growth (of AI capabilities), even in shorter-term horizons. Research Productivity negative AI capabilities growth over time
Reading fidelity high
Study strength medium
not reported
0.06
The METR study claims that fitting sigmoid/logistic curves results in inflection points far in the future. Research Productivity positive timing of inflection point in fitted capability trajectories
Reading fidelity high
Study strength high
not reported
0.1
When we fit a sigmoid curve to METR's current data, the inflection point has already passed. Research Productivity negative timing of inflection point in fitted capability trajectories
Reading fidelity high
Study strength medium
not reported
0.06
We propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, each exhibiting individual rates of improvement. Research Productivity mixed structure of capability growth (base vs reasoning components and their rates)
Reading fidelity high
Study strength speculative
not reported
0.01
We prove that this decomposed model supports the hypothesis that AI capabilities will exhibit an inflection point in the near future. Research Productivity negative presence and timing of an inflection point in AI capability trajectory implied by the model
Reading fidelity high
Study strength speculative
not reported
0.01
The goal of the note is to highlight the fragility of existing forecasts of exponential growth in AI capabilities rather than to establish a rigorous forecast of our own. Research Productivity mixed robustness/fragility of exponential-growth forecasts
Reading fidelity high
Study strength low
not reported
0.03

Notes