4 cumulative citations
View corpus contextFrontier LLM gains are largely a horsepower story—80–90% of top-model advantage maps to more training compute—not hidden proprietary tricks. Away from the frontier, however, developer know‑how and shared algorithmic improvements let some firms build much smaller, more efficient models, and even within a single company model efficiency can vary by over 40×.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Do leading LLM developers possess a proprietary ``secret sauce'', or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, we estimate scaling-law regressions with release-date and developer fixed effects. We find clear evidence of developer-specific efficiency advantages, but their importance depends on where models lie in the performance distribution. At the frontier, 80-90% of performance differences are explained by higher training compute, implying that scale--not proprietary technology--drives frontier advances. Away from the frontier, however, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Some companies can systematically produce smaller models more efficiently. Strikingly, we also find substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. We also discuss the implications for AI leadership and capability diffusion.
Summary
Main Finding
Scaling (training compute) is the dominant driver of frontier LLM performance, but there is robust evidence of a developer “secret sauce” that materially affects performance away from the frontier. Using 809 publicly released LLMs (Oct 2022–Mar 2025) and MMLU‑Pro scores, the authors decompose performance variation into scaling, shared algorithmic progress (time effects), developer-specific effects (“secret sauce”), and model‑specific residuals. At the frontier, 80–90% of performance differences are attributable to larger training compute; away from the frontier, shared and proprietary algorithmic gains substantially reduce the compute needed to reach fixed capability thresholds.
Key Points
- Dataset and outcome
- 809 LLMs released 2022q4–2025q1, benchmarked on MMLU‑Pro (and robustness checks on MATH L5).
- Empirical approach
- Regress logit(MMLU‑Pro) on log(training FLOPs), period dummies, and developer dummies.
- Training compute ci computed as ci = 6 × Ntokens × Nparams (Kaplan et al. approach).
- Use Shapley variance decomposition to apportion explained variance to scaling, shared algorithmic progress (period), company fixed effects, and residual (model) effects.
- Main quantitative results (representative figures)
- Baseline variance shares (all 809 models): scaling ≈ 32%, shared alg. progress ≈ 2.7%, company secret sauce ≈ 14–18% (authors report 14–18% overall), model residuals ≈ 47%.
- Among major developers only: scaling ≈ 45%, shared ≈ 5%, company ≈ 14%, residuals ≈ 36%.
- Frontier models: scaling explains ~80–90% of the rise in top‑score benchmarks (example: Llama2 70B → GPT‑4.5). Training compute at the frontier rose ~5,000× across the sample window.
- Shared algorithmic progress over the window increased effective compute by ≈7.5× (MATH L5 test: ≈11×).
- Developer (company) heterogeneity: some developers show very large estimated compute efficiency factors (e.g., Microsoft ≈60.5× relative to small dev baseline; DeepSeek ≈2.3× — authors note large uncertainty and heterogeneity).
- Within‑company/model heterogeneity: model‑specific residuals imply a 41× difference in effective compute between the 90th and 10th percentile of models (MATH L5 shows even larger dispersion, up to 186×).
- For non‑frontier thresholds: the minimum compute required to reach a 15% normalized MMLU‑Pro score fell ~50× among major developers (≈8,000× if small developers included).
- Regression/fit
- Coefficient on log FLOPs: a 10× increase in training compute raises log‑odds of MMLU‑Pro by ~0.79 (SE 0.19). Model R² reported ~0.52 in the main specification.
- Robustness/auxiliary results
- Similar qualitative patterns for MATH Level 5, with secret sauce relatively more important for math capabilities and larger model‑specific dispersion.
- Results robust to alternative developer sets and period specifications (authors report robustness checks).
Data & Methods
- Outcome variable
- MMLU‑Pro benchmark scores, transformed via logit: Y = ln(y/(1−y)) to match sigmoidal relationship with log compute.
- Key regressors
- log(training compute) — computed from publicly reported parameters and token counts using ci = 6 × Ni × Di.
- Period fixed effects: three roughly equal time blocks (baseline = 2022q4–2023q3; comparisons for later periods capture shared algorithmic improvements).
- Developer fixed effects: dummies for main companies (Deepseek, Qwen, Meta, Google, Microsoft, OpenAI, Anthropic, X‑AI, 01‑AI, Nvidia) plus grouped small/medium/large “other” developers.
- Identification/interpretation
- Period fixed effects interpreted as aggregate/shared algorithmic/technical efficiency gains (diffuse improvements used by many teams).
- Developer fixed effects interpreted as persistent developer‑specific efficiency advantages — the “secret sauce” (captures firm‑level reproducible practices, datasets, fine‑tuning pipelines, teacher models, engineering).
- Model residuals capture model‑specific specializations, fine‑tuning choices, experimentation and measurement noise.
- Decomposition
- OLS estimates feed into a Shapley decomposition of explained variance (R²) to assign shares to each factor.
- Authors also present counterfactual decompositions (e.g., how much of frontier improvement would remain under alternative compute/time/company conditions).
- Limitations noted by authors
- Observational design: fixed effects are descriptive and not causal proofs of mechanisms (e.g., a developer fixed effect bundles many possible channels).
- Measurement: compute is inferred from parameters/tokens with the standard approximation; model performance benchmark choice (MMLU‑Pro) may emphasize some capabilities over others.
- Sample selection: public/reported models only; unpublished internal models and proprietary compute figures are absent.
Implications for AI Economics
- Frontier competition depends on compute access
- Because scale explains most frontier gains, entities that secure superior, expanding compute resources are best positioned to lead in frontier capabilities. Industrial policy and firm strategies that ensure large-scale compute (data‑center capacity, GPUs/accelerators, chilled power, procurement) will be decisive.
- Secret sauce matters for cost‑efficiency and commercial/product competition
- Developer‑level efficiency advantages (14–18% of performance variance overall, higher away from frontier) create persistent cost and product differentiation even when scale is limited. Firms with better data, fine‑tuning, pipelines or engineering can deliver equivalent capabilities at lower compute/inference cost — a commercial advantage for productization and margins.
- Democratization and diffusion below the frontier
- Shared algorithmic improvements and proprietary efficiency gains have drastically lowered the compute required for given capability thresholds (e.g., 15% MMLU‑Pro). This reduces the barriers for smaller developers and countries to build useful/competent models, accelerating diffusion of many LLM capabilities and lowering inference costs for users.
- Risk and governance tradeoffs
- Lower compute requirements for non‑frontier capabilities raise proliferation risks: capable models become easier to train or fine‑tune for misuse. Yet frontier capabilities (most transformative/novel risks) remain concentrated around compute. Governance ought to differentiate policy: controls on large compute supply chains, export controls, or monitoring might target frontier risk, while mitigation of diffusion risks requires monitoring of model releases, datasets, and reuse.
- Strategic firm behavior and market structure
- Two competitive levers exist: (1) invest in scale to pursue frontier leadership; (2) invest in algorithmic/data/process improvements to reduce costs and widen product offerings. Both can coexist: firms may pursue frontier scale for headline capabilities while using secret sauce to improve margins and broaden commercial reach.
- Future of progress if compute growth slows
- If compute availability growth decelerates, the relative importance of algorithmic and firm‑level efficiency will increase. The observed 7.5× shared efficiency gain over ~2 years demonstrates meaningful scope for continued algorithmic returns; however, current data suggest those gains have so far been insufficient to supplant scale at the frontier.
- Research and policy priorities
- Measure and monitor compute concentration and supply chains (because access to compute is central to frontier dominance).
- Encourage transparency on model training (compute, data provenance) to better assess diffusion and risk.
- Support public and academic work on reproducible algorithmic efficiency (to democratize capabilities while giving policymakers better signals).
- Distinguish governance interventions targeted at frontier compute/control versus broader model release practices to address lower‑scale but widespread risks.
Caveats and next steps - The fixed‑effects approach is descriptive; causal inference on which specific practices constitute the “secret sauce” (datasets, fine‑tuning recipes, teacher models, architecture tweaks) requires firm‑level process data or experiments. - The benchmark suite used (MMLU‑Pro, MATH L5) captures particular capability axes; other capabilities may exhibit different scaling/firm effect patterns. - Future work: replicate with private/unreleased models where possible, expand benchmarks, and decompose company effects into measurable components (data quality, prompt‑engineering pipelines, fine‑tuning budgets, etc.).
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We use training and benchmark data for 809 models released between 2022 and 2025. Other | null_result | dataset size / sample of models |
Reading fidelity
high
Study strength
high
|
n=809
|
| We estimate scaling-law regressions with release-date and developer fixed effects. Innovation Output | null_result | relationship between training compute and model performance, controlling for release date and developer |
Reading fidelity
high
Study strength
high
|
n=809
|
| There is clear evidence of developer-specific efficiency advantages. Innovation Output | positive | developer-specific compute efficiency (compute required for given performance) |
Reading fidelity
high
Study strength
medium
|
n=809
|
| The importance of developer-specific advantages depends on where models lie in the performance distribution. Innovation Output | mixed | heterogeneity of developer efficiency across performance quantiles |
Reading fidelity
high
Study strength
medium
|
n=809
|
| At the frontier, 80–90% of performance differences are explained by higher training compute, implying that scale — not proprietary technology — drives frontier advances. Innovation Output | positive | share of performance differences explained by training compute at frontier models |
Reading fidelity
high
Study strength
medium
|
n=809
80-90% of performance differences explained by higher training compute
|
| Away from the frontier, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Innovation Output | positive | compute required to reach fixed capability thresholds outside the frontier |
Reading fidelity
high
Study strength
medium
|
n=809
|
| Some companies can systematically produce smaller models more efficiently. Innovation Output | positive | compute-to-capability efficiency of specific companies |
Reading fidelity
high
Study strength
medium
|
n=809
|
| There is substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. Innovation Output | mixed | within-firm compute-efficiency variation (ratio of compute required for models achieving similar capabilities) |
Reading fidelity
high
Study strength
medium
|
n=809
more than 40x compute efficiency difference
|
| We discuss the implications for AI leadership and capability diffusion. Governance And Regulation | null_result | implications for AI leadership and capability diffusion (qualitative) |
Reading fidelity
high
Study strength
speculative
|
n=809
|