Quality-adjusted prices for AI inference services have fallen roughly seven times faster than standard matched-model measures indicate, implying most price declines are hidden from current economic statistics; when measured per completed task, buyer-facing prices have stopped falling.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurement. This paper constructs quality-adjusted price indices for the AI inference market from public data. The panel assembles 21,024 posted-price observations across 3,208 models and 86 providers and joins them to 4,605 benchmark scores through a latent quality index estimated from benchmark response patterns, so the quality ladder of the hedonic tradition is built here from evaluations in place of product characteristics. Measured by the matched-model methods that statistical agencies apply to software, inference prices fell at 0.10 log points a year. The quality-adjusted index fell at 0.73, so 87% of the decline is invisible to current methods, with direct consequences for measured competition, concentration and productivity in this market. Counted per completed task, moreover, the buyer's price stopped falling. Reasoning models raised token consumption faster than token prices fell, and the seller's and buyer's prices accordingly diverged. A pre-registered validity audit disciplines the quality measure and yields the sharpest result. Excluding contamination-flagged benchmarks leaves model rankings intact at 0.998 yet moves the index by 0.49 log points a year, so the leaderboard-stability arguments standard in AI evaluation offer no defence of economic statistics built on benchmarks. Prices, quality and the audit are fully reproducible from public sources at zero cost.
Summary
Main Finding
Quality-adjusting posted prices for AI inference using benchmark-based capability estimates shows that most of the apparent (list) price decline is driven by quality improvements. Measured as statistical agencies currently do (matched-model methods), posted inference prices fell ≈ −0.10 log points/year (2025–2026). After hedonic quality adjustment using an equated capability index, prices fell ≈ −0.73 log points/year — meaning roughly 87% of the measured decline is invisible to current methods. Measured per completed task (buyer unit) the price stopped falling over the observable window because token consumption rose faster than token prices fell.
Key Points
- Data scope: 21,024 posted-price observations across 3,208 priced models, 86 providers, 31 months (Feb 2024–Aug 2026); 4,605 scored model–benchmark cells across 64 benchmarks; dataset and code publicly released (CC BY / MIT) and fully reproducible.
- Capability estimation: continuous-response latent-variable model (Samejima-style 2PL analogue for continuous scores) fit to the score matrix; each model receives θm with posterior SEs (median SE 0.135; range 0.033–3.337). Residual SD shared and estimated at 0.449.
- Equating: two-stage fixed-parameter equating across benchmark generations, with anchor selection rule (≥10 scored cells each side of 1 July 2024 and score SD ≥ 0.10). GPQA Diamond was pinned for scale (a=1, d=0) as an identification convention.
- Uncertainty: capability uncertainty propagated downstream by drawing L = 20 plausible θ values per model and combining results via Rubin’s rules.
- Price indices: rolling-window hedonic (time-dummy) index and matched-model alternatives estimated; indices reported in two units — per-token (seller unit) and per-completed-task (buyer unit).
- Main numeric comparisons: matched-model (statistical-agency style) posted-price decline ≈ −0.10 log points/year; quality-adjusted (benchmark hedonic) posted-price decline ≈ −0.73 log points/year; exclusion of contamination-flagged benchmarks shifts the index by ≈ 0.49 log points/year.
- Buyer vs seller divergence: reasoning models increased token consumption per task faster than token prices fell, producing divergence such that buyer price per completed task ceased declining even as per-token list prices fell.
- Audit (pre-registered, argument-based following Kane): four inferences — Scoring, Generalisation, Extrapolation, Decision. Two of four fail in material ways:
- Scoring inference: many anchor benchmarks showed differential functioning consistent with contamination; six of nine testable anchors flagged; for seven of ten anchors the standard contamination test had no suitable comparison group (anchors predate most models).
- Excluding benchmarks flagged for contamination preserves model ranking (Spearman correlation ≈ 0.998) but materially re-scales capability and moves the price index by ≈ 0.49 log points/year. Thus ranking robustness is not a substitute for validity of scale-spacing.
- Decision inference “passes” but with reversed sign relative to expectation: capability measurement quality is worse in the sparsely evaluated back-catalogue and best at the frontier (evaluation attention concentrates on the frontier), implying bias concerns run opposite to naive intuition.
- Data provenance: benchmark scores and metadata from Epoch AI Benchmark Hub; posted prices and token consumption from Artificial Analysis and Internet Archive; cross-checks from OpenRouter, LMArena, and provider pages. Every price observation keeps source URL and capture time; a 5% random sample of observed rows (1,051 rows) released for inspection.
- Important caveat: these are posted list prices on first-party channels; transaction prices (after discounts, enterprise deals, batching, caching, etc.) are not public and are likely lower and declining faster — so the paper likely understates true buyer price declines.
Data & Methods
- Data assembled monthly (archived pages provide embedded structured state enabling recovery of historical posted prices that rendered tables do not show).
- Price panel: 21,024 observed rows, 3,208 priced models, 86 providers, Feb 2024–Aug 2026. Token-per-task consumption available from mid-2026 source; task-denominated index covers a shorter window.
- Capability model:
- Observations: 4,605 model–benchmark score cells over 64 benchmarks and 782 models used in the scoring model.
- Model: logit(s_mk) = a_k (θ_m − d_k) + ε_mk, ε ~ N(0, σ^2), continuous-response/IRT-style; boundary scores nudged inward by ε = 0.001.
- Equating: fixed-anchor, two-stage estimation to link benchmark generations; anchors chosen by pre-registered rule (≥10 cells per era and SD ≥0.10); ten anchors selected.
- Identification: fixes anchor benchmark parameters (e.g., GPQA Diamond at a=1, d=0) to set scale.
- Uncertainty: 20 posterior draws per θ, all downstream computations (hedonic regressions, indices) recomputed per draw; intervals combined with Rubin’s rules.
- Price indices:
- Hedonic rolling-window time-dummy (splicing) estimator for quality-adjusted index, reflecting standard agency practice but substituting θ for physical characteristics.
- Matched-model methods (statistical-agency style) estimated for comparison.
- Indices reported in seller unit (per-token) and buyer unit (per completed task). Token consumption data limit task-index window.
- Auditing protocol:
- Pre-registered validity audit (Kane-style chain of inferences): Scoring (differential item functioning / contamination), Generalisation (out-of-sample performance), Extrapolation (ability outside tests), Decision (index-level suitability).
- Differential functioning tests applied to anchor benchmarks; contamination flags based on unsupported functioning patterns and provenance (many early anchors are developer- or self-reported and predate most models).
- Robustness and checks:
- External comparisons: θ correlates strongly with release year (Spearman 0.880) and partially agrees with item-level IRT in benchmarks where per-question logs are available (3/5 benchmarks ≳ 0.7 rank correlation).
- Reproducibility: all inputs and code embedded with cryptographic hashes; public repository with pre-registration and deviation log.
Implications for AI Economics
- Measurement: current indexation/deflation methods (matched-models without benchmark quality adjustment) substantially understate the real quality-adjusted decline in inference prices — this affects GDP deflators, productivity accounting, and real output measurement for AI services.
- Market analysis: understated price decline leads to biased inferences on competition, concentration, and market power. If quality gains are the main driver, measured price stagnation might mask intense competition via rapid capability improvements at the same list prices.
- Policy and statistical practice:
- Benchmark-based quality adjustment is feasible and large in magnitude; statistical agencies should consider adopting it — but only with explicit, adversarial validity auditing of benchmarks (structure, contamination, provenance).
- Ranking robustness (leaderboard stability) is insufficient; scale spacing matters for hedonic adjustment. Validation must interrogate metric structure (differential functioning, contamination) and re-anchor if needed.
- Indices must propagate capability uncertainty into final intervals; otherwise inferential statements will be overconfident.
- Units matter: choice of unit (per-token seller vs per-task buyer) can flip the sign of measured real growth. Analysts and policymakers must be explicit about the economic object of interest (tokens vs completed tasks).
- Data gaps: public posted prices are imperfect proxies for transaction prices (discounts, enterprise deals). Availability of transaction-level pricing would improve accuracy but is currently absent; the reported results should be read as conservative bounds on true declines.
- Evaluation infrastructure: the current evaluation ecosystem cannot automatically certify that benchmarks are uncontaminated or suitable as hedonic characteristics. Independent auditing and provenance tracking should be part of any pipeline that feeds statistical measurement.
- Practical recommendations implied by the paper:
- Adopt benchmark-based hedonic adjustments but accompany them with pre-registered validity audits and sensitivity analyses (e.g., exclusion of flagged anchors).
- Require provenance metadata and archive score tables; prefer benchmarks with linking power across eras and demonstrable non-differential functioning.
- Report both seller- and buyer-unit indices and propagate measurement uncertainty to policy users.
Limitations and cautions: anchor choices and fixed-anchor equating can propagate anchor defects; posted list-price basis understates true price declines; token-consumption series are shorter and back-measurement assumptions matter for the task-denominated index. The paper’s dataset and methodology are public, enabling replication and alternative specifications.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Measured using matched-model methods applied to software, AI inference prices fell by 0.10 log points per year from 2025 to 2026. Firm Productivity | negative | Matched-model posted price change for AI inference |
Reading fidelity
high
Study strength
medium
|
n=21024
0.10 log points a year
|
| The quality-adjusted AI inference price index fell by 0.73 log points per year. Firm Productivity | negative | Quality-adjusted posted price change for AI inference |
Reading fidelity
high
Study strength
medium
|
n=21024
0.73 log points a year
|
| Quality improvement accounts for 87% of the decline measured by the quality-adjusted index that is not visible in matched-model methods. Firm Productivity | negative | Share of AI inference price decline attributable to quality adjustment |
Reading fidelity
high
Study strength
medium
|
n=21024
87% of the decline
|
| Counted per completed task, the buyer's price stopped falling over the period for which task-level measurement is available. Consumer Welfare | null_result | Buyer price per completed AI task |
Reading fidelity
high
Study strength
medium
|
n=20
|
| Reasoning models increased token consumption faster than token prices declined, causing seller prices per token and buyer prices per completed task to diverge. Consumer Welfare | mixed | Relationship between token prices and task-level buyer prices |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Excluding contamination-flagged benchmarks leaves model rankings nearly unchanged, with a correlation of 0.998, but changes the estimated quality-adjusted price decline by 0.49 log points per year. Firm Productivity | mixed | Sensitivity of model capability rankings and the quality-adjusted price index to benchmark contamination exclusions |
Reading fidelity
high
Study strength
medium
|
n=4605
correlation of 0.998; 0.49 log points a year
|
| The study's frozen panel contains 21,024 observed price rows for 3,208 priced models from 86 providers over 31 months, alongside 4,605 model-benchmark score cells across 64 benchmarks. Other | other | Scope of the AI inference price and benchmark panel |
Reading fidelity
high
Study strength
medium
|
n=21024
|
| All 40 price observations in the hand-checked accuracy sample matched their recorded source URLs without correction. Other | positive | Accuracy of recorded posted-price observations |
Reading fidelity
high
Study strength
low
|
n=40
40 matching without correction
|
| The latent capability measure was estimated from 4,605 scored cells covering 782 models and 64 benchmarks. Other | positive | Estimated AI model capability |
Reading fidelity
high
Study strength
medium
|
n=4605
|
| Estimated capability is strongly associated with model release date, with a Spearman correlation of 0.880 across 757 dated models. Other | positive | Association between estimated capability and model release date |
Reading fidelity
high
Study strength
low
|
n=757
Spearman 0.880
|
| Six of nine testable benchmark anchors show differential functioning consistent with contamination, and seven of ten anchors lack a comparison group for the standard contamination test. Ai Safety And Ethics | negative | Benchmark validity and contamination detectability |
Reading fidelity
high
Study strength
medium
|
n=10
6 of 9; 7 of 10
|