0 cumulative citations
View corpus contextAn audit finds frontier-AI progress data is fragmentary and source-concentrated, so fitting a date from a single score or compute trend is often unsupported; robust forecasts require explicit, versioned measurement systems with designed linking and provenance transparency.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.
Summary
Main Finding
Frontier-AI forecasting faces a fundamental measurement problem: the public longitudinal record needed to convert benchmark scores, training compute, and other indicators into defensible dated forecasts is fragmented and instrument-dependent. Sparse joint observability, unstable/ambiguous benchmark linking, and concentrated provenance (few sources producing most observations) make simple extrapolations (fit a curve → read off a date) unreliable unless the forecast explicitly conditions on a versioned measurement system (joins, linking functions, protocols and source dependence).
Key Points
- Audit scope and snapshot
- Evidence frozen at 12 Aug 2026.
- Analytic sample: 62 selected systems (27 closed, 35 open-weight), 12 benchmark-version records, 7 capability/impact criteria, 144 graded events, 27 source identifiers, 408 typed relations.
- Audit is an event-centric, reproducible evidence ledger (not a census).
- Sparse joint observability
- Training-compute values recorded for 43/62 systems, METR p50/p80 for 26 systems.
- Only 7 of 62 systems jointly observe training compute and METR p50 — the key join for a compute→capability progression analysis.
- Training compute is missing for 19/27 closed systems (including all selected closed releases from 2026). No open-weight system in the sample has a METR horizon observation.
- Missingness correlates with access regime (closed vs open-weight) and thus is unlikely ignorable.
- Benchmark versioning and linking problems
- Benchmarks are versioned instruments; protocol changes, corrections, contamination, and measurement ceilings are frequent and matter for longitudinal inference.
- Example bridges:
- METR Time Horizon 1.0 → 1.1 (7 common systems): fitted log2 slope = 1.206 (95% CI 1.021–1.390) — not well represented by a pure additive shift on log scale (shape change evident).
- MMLU → MMLU-Pro (6 common systems): slopes vary by score transform (logit slope ≈ 0.976; probit ≈ 1.043; linear ≈ 1.312; log2 ≈ 1.984) — diagnosis depends on chosen link.
- Link estimates and conclusions are scale-dependent; analysts must specify measurement model and linking design rather than concatenating versions.
- Provenance concentration and replication limits
- Of 71 substantive quantitative events (benchmark results, field experiments, substrate estimates, efficiency observations), 52 (73.2%) come from a single programme (METR Time Horizon 1.1).
- 76.1% of substantive events are laboratory releases; only 6 peer-reviewed events.
- Source Herfindahl index = 0.542 (≈1.85 effective equally represented sources): many observations are not independent replications but share programme-specific protocols and potential correlated errors.
- Power and detectability
- Observed linking designs have ~80% power only for slope departures near ~25%.
- METR bridge: observed power ≈ 63% for the observed slope departure.
- MMLU logit bridge: observed power ≈ 6%.
- To detect a 10% slope departure (under observed noise/spread) would require roughly 23–31 common systems (anchor panels spanning score range).
- Measurement portfolio perspective
- No single scalar replacement exists. The audit identifies 16 complementary measurement directions (resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, backtests, etc.).
- Multiple design requirements are necessary: anchored linking panels, psychometric/item-response models, reporting of resource and inference budgets, reliability curves, field validation, and provenance/replication metadata.
Data & Methods
- Evidence freeze and selection
- Cutoff date: 12 Aug 2026. Sample chosen to support compute–capability audit and include historically important public releases; not exhaustive.
- Training-compute estimates drawn from Epoch AI and processed Our World in Data series; METR Time Horizon observations frozen locally.
- Event-centric representation
- Normalized, versioned entities for models, benchmark versions, events and sources; event table timestamps releases, benchmark scores, revisions, contamination/correction events, field experiments, substrate/efficiency estimates.
- Graph contains 269 nodes and 408 typed relations; every substantive quantitative event has a source id and an evidence grade.
- Empirical diagnostics
- Joint-observability matrix recorded per model: presence/absence of training compute, METR p50/p80, any capability result, non-METR capability result, peer-reviewed measurement.
- Benchmark linking stress tests: fit g(Ynew) = α + β g(Yold) + ε using different transforms (log2 for time horizons; logit/probit/linear/log for bounded accuracy measures). OLS used as a transparent diagnostic; leave-one-out sensitivity reported.
- Power calculations: minimum detectable effect (|β−1|) computed for two-sided tests at α = 0.05, 80% power using the noncentral t distribution; projections hold observed residual noise and anchor-score dispersion fixed (design illustrations).
- Source & literature audit
- Administrative release rows excluded from provenance-concentration stats.
- Coded 56 external sources (29 peer-reviewed, 27 preprints/standards/reports/datasets) into 16 measurement directions and assessed design maturity (qualitative scoring 0–3).
- Documented nine recurring evidence-design requirements emerging from the literature and audit.
Implications for AI Economics
- Fragility of timeline-based economic forecasts
- Many economic models of AI impacts (labor substitution, productivity growth, capital reallocation, investment timing, insurance and policy decisions) rely on dated capability thresholds. Given the sparse and non-random observability of compute–capability joins and unstable benchmark links, timeline forecasts without explicit measurement-system conditioning risk giving false precision.
- Need to propagate measurement uncertainty into economic analyses
- Economists should treat dated capability forecasts as claims about a measurement system (instrument version, linking function, resource/inference assumptions, and source dependence), and propagate uncertainty arising from instrument changes and provenance concentration into downstream economic predictions (e.g., timing of automation, value of retraining, investment horizons).
- Design requirements for forecastable indicators
- For usable inputs to economic models, evaluation programs should:
- Produce joint observations of predictors (training compute, data, parameters) and outcome measures across diverse developer ecosystems.
- Implement designed anchor/linking panels (multiple common systems/items spanning score ranges) when benchmarks are revised.
- Adopt psychometric/item-response methods and report uncertainty, judge variability and calibration across versions.
- Report inference budgets, scaffolding, prompting protocols and resource costs (price-performance, energy) to connect capabilities to deployment costs.
- Increase replication across independent programmes (reduce provenance concentration) and publish provenance metadata.
- For usable inputs to economic models, evaluation programs should:
- Implications for policy and investment
- Fund and support interoperable, versioned evaluation infrastructure (anchor datasets, shared testbeds, standard provenance metadata) to improve forecastability.
- Encourage disclosure norms (or standardized reporting) around training compute and training data summaries to reduce missingness bias that correlates with access regime.
- Prioritize field validation (real-world task performance) and backtested forecasts to build evidence on how lab scores translate to economic outcomes.
- Practical advice for economic modelers
- Avoid treating a single benchmark timeseries as a latent scalar trend; instead model capability progress conditional on explicit measurement instruments and include measurement-model uncertainty.
- When using compute-based progression axes, require or impute joint observations with careful sensitivity checks to disclosure patterns (closed vs open releases).
- Where possible, combine multiple complementary measures (resource frontiers, reliability curves, agentic task performance, human preference/utility measures) rather than relying on a single score to infer economic impact timing.
Summary takeaway: Forecasts that produce precise dates for "frontier AI" milestones should be read as conditional on specific, versioned measurement systems. Improving the reliability of AI-economics forecasts requires investment in longitudinal measurement design (anchors, linking, provenance, replication) and explicit propagation of measurement uncertainty into economic models.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The audit record contains 62 selected AI systems, consisting of 27 closed systems and 35 open-weight systems; it is an audit sample rather than a census of releases. Automation Exposure | null_result | Coverage and composition of the frontier-AI measurement record |
Reading fidelity
high
Study strength
medium
|
n=62
|
| Training compute is available for 43 of 62 systems, but 19 of 27 closed systems lack a compute estimate, while all 35 open-weight systems in the sample have one. Automation Exposure | negative | Availability of training-compute observations |
Reading fidelity
high
Study strength
medium
|
n=62
19 of 27 closed systems lack compute estimates; 35 of 35 open-weight systems have estimates
|
| Only seven systems jointly observe estimated training compute and METR p50 task horizon, and no open-weight system in the sample provides the intended compute–horizon join. Automation Exposure | negative | Joint observability of training compute and task horizon |
Reading fidelity
high
Study strength
medium
|
n=62
7 of 62 systems
|
| The apparent compute–task-horizon relationship in the audit is identified by only seven historically clustered systems and is descriptive rather than evidence of a stable scaling law or causal effect. Task Completion Time | null_result | Association between estimated training compute and METR p50 task horizon |
Reading fidelity
high
Study strength
low
|
n=7
descriptive log–log slope 0.73
|
| Measurement provenance is highly concentrated: 52 of 71 substantive quantitative events, or 73.2%, come from the METR Time Horizon 1.1 programme. Research Productivity | negative | Concentration and independence of quantitative measurement sources |
Reading fidelity
high
Study strength
medium
|
n=71
73.2%
|
| The 71 substantive quantitative events are predominantly laboratory releases: 54 events, or 76.1%, are laboratory releases. Research Productivity | negative | Distribution of evidence by source or venue type |
Reading fidelity
high
Study strength
medium
|
n=71
76.1%
|
| For seven systems measured on METR Time Horizon 1.0 and 1.1, the fitted log2 linking slope is 1.206, with a 95% confidence interval of 1.021–1.390, indicating that a pure additive shift on the log scale does not fit the bridge well. Task Completion Time | positive | Change in METR task-horizon measurements across benchmark versions |
Reading fidelity
high
Study strength
medium
|
n=7
log2 slope 1.206 (95% CI 1.021–1.390)
|
| The apparent relationship between MMLU and MMLU-Pro scores depends strongly on the chosen scale: for six common systems, the fitted slope is 0.976 under a logit link, 1.043 under a probit link, 1.312 on the linear scale, and 1.984 on the logarithmic scale. Output Quality | mixed | Cross-version or cross-benchmark score comparability |
Reading fidelity
high
Study strength
low
|
n=6
logit slope 0.976; probit slope 1.043; linear slope 1.312; logarithmic slope 1.984
|
| The observed benchmark-linking designs have roughly 80% power only for slope departures of about 25% from one; the METR bridge has 63% power at its observed departure, while the MMLU logit bridge has 6% power. Decision Quality | negative | Statistical power to detect benchmark-linking slope departures |
Reading fidelity
high
Study strength
medium
|
n=13
80% power near 25% slope departures; 63% observed power for METR; 6% observed power for MMLU logit
|
| Under the audit's same-noise and same-score-spread assumptions, detecting a 10% departure from a shift-like linking slope would require approximately 23–31 common systems. Decision Quality | negative | Required sample size for detecting benchmark-linking departures |
Reading fidelity
high
Study strength
medium
|
n=31
approximately 23–31 common systems for a 10% departure
|
| The review of 56 additional methodological and empirical sources identified 16 complementary measurement directions, but no direction supplied a replacement universal scalar for frontier-AI progress. Other | null_result | Availability of a universal scalar measure of frontier-AI progress |
Reading fidelity
high
Study strength
low
|
n=56
16 measurement directions; no replacement scalar
|