The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Quality-adjusted prices for AI inference services have fallen roughly seven times faster than standard matched-model measures indicate, implying most price declines are hidden from current economic statistics; when measured per completed task, buyer-facing prices have stopped falling.

The Price of Intelligence: A Quality-Adjusted Price Index for AI Services
Louis Yiven Zhu · August 30, 2026
arxiv descriptive high evidence 9/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Louis Yiven Zhu unresolved corpus identity

Semantic Scholar

Latest observation:

  1. L. Zhu provider ID
Using a reproducible public panel and a latent capability index, the paper finds quality-adjusted AI inference prices fell about 0.73 log points per year versus 0.10 by matched-model methods—implying 87% of the apparent decline is invisible to current statistical practice—while per-task buyer prices stopped falling.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurement. This paper constructs quality-adjusted price indices for the AI inference market from public data. The panel assembles 21,024 posted-price observations across 3,208 models and 86 providers and joins them to 4,605 benchmark scores through a latent quality index estimated from benchmark response patterns, so the quality ladder of the hedonic tradition is built here from evaluations in place of product characteristics. Measured by the matched-model methods that statistical agencies apply to software, inference prices fell at 0.10 log points a year. The quality-adjusted index fell at 0.73, so 87% of the decline is invisible to current methods, with direct consequences for measured competition, concentration and productivity in this market. Counted per completed task, moreover, the buyer's price stopped falling. Reasoning models raised token consumption faster than token prices fell, and the seller's and buyer's prices accordingly diverged. A pre-registered validity audit disciplines the quality measure and yields the sharpest result. Excluding contamination-flagged benchmarks leaves model rankings intact at 0.998 yet moves the index by 0.49 log points a year, so the leaderboard-stability arguments standard in AI evaluation offer no defence of economic statistics built on benchmarks. Prices, quality and the audit are fully reproducible from public sources at zero cost.

Summary

Main Finding

Quality-adjusting posted prices for AI inference using benchmark-based capability estimates shows that most of the apparent (list) price decline is driven by quality improvements. Measured as statistical agencies currently do (matched-model methods), posted inference prices fell ≈ −0.10 log points/year (2025–2026). After hedonic quality adjustment using an equated capability index, prices fell ≈ −0.73 log points/year — meaning roughly 87% of the measured decline is invisible to current methods. Measured per completed task (buyer unit) the price stopped falling over the observable window because token consumption rose faster than token prices fell.

Key Points

  • Data scope: 21,024 posted-price observations across 3,208 priced models, 86 providers, 31 months (Feb 2024–Aug 2026); 4,605 scored model–benchmark cells across 64 benchmarks; dataset and code publicly released (CC BY / MIT) and fully reproducible.
  • Capability estimation: continuous-response latent-variable model (Samejima-style 2PL analogue for continuous scores) fit to the score matrix; each model receives θm with posterior SEs (median SE 0.135; range 0.033–3.337). Residual SD shared and estimated at 0.449.
  • Equating: two-stage fixed-parameter equating across benchmark generations, with anchor selection rule (≥10 scored cells each side of 1 July 2024 and score SD ≥ 0.10). GPQA Diamond was pinned for scale (a=1, d=0) as an identification convention.
  • Uncertainty: capability uncertainty propagated downstream by drawing L = 20 plausible θ values per model and combining results via Rubin’s rules.
  • Price indices: rolling-window hedonic (time-dummy) index and matched-model alternatives estimated; indices reported in two units — per-token (seller unit) and per-completed-task (buyer unit).
  • Main numeric comparisons: matched-model (statistical-agency style) posted-price decline ≈ −0.10 log points/year; quality-adjusted (benchmark hedonic) posted-price decline ≈ −0.73 log points/year; exclusion of contamination-flagged benchmarks shifts the index by ≈ 0.49 log points/year.
  • Buyer vs seller divergence: reasoning models increased token consumption per task faster than token prices fell, producing divergence such that buyer price per completed task ceased declining even as per-token list prices fell.
  • Audit (pre-registered, argument-based following Kane): four inferences — Scoring, Generalisation, Extrapolation, Decision. Two of four fail in material ways:
    • Scoring inference: many anchor benchmarks showed differential functioning consistent with contamination; six of nine testable anchors flagged; for seven of ten anchors the standard contamination test had no suitable comparison group (anchors predate most models).
    • Excluding benchmarks flagged for contamination preserves model ranking (Spearman correlation ≈ 0.998) but materially re-scales capability and moves the price index by ≈ 0.49 log points/year. Thus ranking robustness is not a substitute for validity of scale-spacing.
    • Decision inference “passes” but with reversed sign relative to expectation: capability measurement quality is worse in the sparsely evaluated back-catalogue and best at the frontier (evaluation attention concentrates on the frontier), implying bias concerns run opposite to naive intuition.
  • Data provenance: benchmark scores and metadata from Epoch AI Benchmark Hub; posted prices and token consumption from Artificial Analysis and Internet Archive; cross-checks from OpenRouter, LMArena, and provider pages. Every price observation keeps source URL and capture time; a 5% random sample of observed rows (1,051 rows) released for inspection.
  • Important caveat: these are posted list prices on first-party channels; transaction prices (after discounts, enterprise deals, batching, caching, etc.) are not public and are likely lower and declining faster — so the paper likely understates true buyer price declines.

Data & Methods

  • Data assembled monthly (archived pages provide embedded structured state enabling recovery of historical posted prices that rendered tables do not show).
  • Price panel: 21,024 observed rows, 3,208 priced models, 86 providers, Feb 2024–Aug 2026. Token-per-task consumption available from mid-2026 source; task-denominated index covers a shorter window.
  • Capability model:
    • Observations: 4,605 model–benchmark score cells over 64 benchmarks and 782 models used in the scoring model.
    • Model: logit(s_mk) = a_k (θ_m − d_k) + ε_mk, ε ~ N(0, σ^2), continuous-response/IRT-style; boundary scores nudged inward by ε = 0.001.
    • Equating: fixed-anchor, two-stage estimation to link benchmark generations; anchors chosen by pre-registered rule (≥10 cells per era and SD ≥0.10); ten anchors selected.
    • Identification: fixes anchor benchmark parameters (e.g., GPQA Diamond at a=1, d=0) to set scale.
    • Uncertainty: 20 posterior draws per θ, all downstream computations (hedonic regressions, indices) recomputed per draw; intervals combined with Rubin’s rules.
  • Price indices:
    • Hedonic rolling-window time-dummy (splicing) estimator for quality-adjusted index, reflecting standard agency practice but substituting θ for physical characteristics.
    • Matched-model methods (statistical-agency style) estimated for comparison.
    • Indices reported in seller unit (per-token) and buyer unit (per completed task). Token consumption data limit task-index window.
  • Auditing protocol:
    • Pre-registered validity audit (Kane-style chain of inferences): Scoring (differential item functioning / contamination), Generalisation (out-of-sample performance), Extrapolation (ability outside tests), Decision (index-level suitability).
    • Differential functioning tests applied to anchor benchmarks; contamination flags based on unsupported functioning patterns and provenance (many early anchors are developer- or self-reported and predate most models).
  • Robustness and checks:
    • External comparisons: θ correlates strongly with release year (Spearman 0.880) and partially agrees with item-level IRT in benchmarks where per-question logs are available (3/5 benchmarks ≳ 0.7 rank correlation).
    • Reproducibility: all inputs and code embedded with cryptographic hashes; public repository with pre-registration and deviation log.

Implications for AI Economics

  • Measurement: current indexation/deflation methods (matched-models without benchmark quality adjustment) substantially understate the real quality-adjusted decline in inference prices — this affects GDP deflators, productivity accounting, and real output measurement for AI services.
  • Market analysis: understated price decline leads to biased inferences on competition, concentration, and market power. If quality gains are the main driver, measured price stagnation might mask intense competition via rapid capability improvements at the same list prices.
  • Policy and statistical practice:
    • Benchmark-based quality adjustment is feasible and large in magnitude; statistical agencies should consider adopting it — but only with explicit, adversarial validity auditing of benchmarks (structure, contamination, provenance).
    • Ranking robustness (leaderboard stability) is insufficient; scale spacing matters for hedonic adjustment. Validation must interrogate metric structure (differential functioning, contamination) and re-anchor if needed.
    • Indices must propagate capability uncertainty into final intervals; otherwise inferential statements will be overconfident.
  • Units matter: choice of unit (per-token seller vs per-task buyer) can flip the sign of measured real growth. Analysts and policymakers must be explicit about the economic object of interest (tokens vs completed tasks).
  • Data gaps: public posted prices are imperfect proxies for transaction prices (discounts, enterprise deals). Availability of transaction-level pricing would improve accuracy but is currently absent; the reported results should be read as conservative bounds on true declines.
  • Evaluation infrastructure: the current evaluation ecosystem cannot automatically certify that benchmarks are uncontaminated or suitable as hedonic characteristics. Independent auditing and provenance tracking should be part of any pipeline that feeds statistical measurement.
  • Practical recommendations implied by the paper:
    • Adopt benchmark-based hedonic adjustments but accompany them with pre-registered validity audits and sensitivity analyses (e.g., exclusion of flagged anchors).
    • Require provenance metadata and archive score tables; prefer benchmarks with linking power across eras and demonstrable non-differential functioning.
    • Report both seller- and buyer-unit indices and propagate measurement uncertainty to policy users.

Limitations and cautions: anchor choices and fixed-anchor equating can propagate anchor defects; posted list-price basis understates true price declines; token-consumption series are shorter and back-measurement assumptions matter for the task-denominated index. The paper’s dataset and methodology are public, enabling replication and alternative specifications.

Assessment

Paper Typedescriptive Evidence Strengthhigh — The paper assembles a large, fully reproducible public panel, pre-registers its analysis, propagates measurement uncertainty, and runs a structured validity audit; these practices give strong empirical support for its measurement claims, though important caveats (posted list vs transaction prices, benchmark contamination/saturation) are transparently discussed and affect magnitude estimates. Methods Rigorhigh — Uses a principled latent-variable (continuous response) model to equate heterogeneous benchmarks, pins anchors with a pre-specified rule, propagates posterior draws through all downstream statistics, pre-registers analyses, and conducts a multi-step validity audit; some methodological choices (fixed-anchor equating, shared residual variance, reliance on posted prices and public benchmarks) introduce identifiable sensitivities which the paper tests but do not eliminate. SamplePanel of 21,024 posted-list price observations across 3,208 priced models, 86 providers and 31 months (Feb 2024–Aug 2026); 4,605 scored model–benchmark cells covering 782 models and 64 benchmarks (3,796 distinct model–benchmark pairs); per-task token consumption for a subset (task index covers a shorter window). Sources include Epoch AI Benchmark Hub, Artificial Analysis plus Internet Archive captures, OpenRouter, LMArena and provider pages; every price observation records URL and timestamp and a 5% random sample of rows is released for inspection. Themesproductivity governance GeneralizabilityUses posted list prices (per-channel) rather than transaction prices, so measured declines likely understate true consumer price declines, Coverage limited to publicly observable models, providers and benchmarks; private enterprise agreements, committed discounts and unobserved channels are excluded, Benchmark-based quality adjustment depends on validity of public benchmarks; contamination, saturation and heterogeneous provenance limit inference, Short sample window (2024–2026) captures a specific early phase of the market and may not generalize to later equilibria, Index weighting lacks expenditure/market-share weights (no transaction volumes), limiting interpretation for aggregate welfare

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Measured using matched-model methods applied to software, AI inference prices fell by 0.10 log points per year from 2025 to 2026. Firm Productivity negative Matched-model posted price change for AI inference
Reading fidelity high
Study strength medium
n=21024
0.10 log points a year
0.18
The quality-adjusted AI inference price index fell by 0.73 log points per year. Firm Productivity negative Quality-adjusted posted price change for AI inference
Reading fidelity high
Study strength medium
n=21024
0.73 log points a year
0.18
Quality improvement accounts for 87% of the decline measured by the quality-adjusted index that is not visible in matched-model methods. Firm Productivity negative Share of AI inference price decline attributable to quality adjustment
Reading fidelity high
Study strength medium
n=21024
87% of the decline
0.18
Counted per completed task, the buyer's price stopped falling over the period for which task-level measurement is available. Consumer Welfare null_result Buyer price per completed AI task
Reading fidelity high
Study strength medium
n=20
0.18
Reasoning models increased token consumption faster than token prices declined, causing seller prices per token and buyer prices per completed task to diverge. Consumer Welfare mixed Relationship between token prices and task-level buyer prices
Reading fidelity high
Study strength medium
not reported
0.18
Excluding contamination-flagged benchmarks leaves model rankings nearly unchanged, with a correlation of 0.998, but changes the estimated quality-adjusted price decline by 0.49 log points per year. Firm Productivity mixed Sensitivity of model capability rankings and the quality-adjusted price index to benchmark contamination exclusions
Reading fidelity high
Study strength medium
n=4605
correlation of 0.998; 0.49 log points a year
0.18
The study's frozen panel contains 21,024 observed price rows for 3,208 priced models from 86 providers over 31 months, alongside 4,605 model-benchmark score cells across 64 benchmarks. Other other Scope of the AI inference price and benchmark panel
Reading fidelity high
Study strength medium
n=21024
0.18
All 40 price observations in the hand-checked accuracy sample matched their recorded source URLs without correction. Other positive Accuracy of recorded posted-price observations
Reading fidelity high
Study strength low
n=40
40 matching without correction
0.09
The latent capability measure was estimated from 4,605 scored cells covering 782 models and 64 benchmarks. Other positive Estimated AI model capability
Reading fidelity high
Study strength medium
n=4605
0.18
Estimated capability is strongly associated with model release date, with a Spearman correlation of 0.880 across 757 dated models. Other positive Association between estimated capability and model release date
Reading fidelity high
Study strength low
n=757
Spearman 0.880
0.09
Six of nine testable benchmark anchors show differential functioning consistent with contamination, and seven of ten anchors lack a comparison group for the standard contamination test. Ai Safety And Ethics negative Benchmark validity and contamination detectability
Reading fidelity high
Study strength medium
n=10
6 of 9; 7 of 10
0.18

Notes