The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Frontier LLM gains are largely a horsepower story—80–90% of top-model advantage maps to more training compute—not hidden proprietary tricks. Away from the frontier, however, developer know‑how and shared algorithmic improvements let some firms build much smaller, more efficient models, and even within a single company model efficiency can vary by over 40×.

Is there "Secret Sauce'' in Large Language Model Development?
Matthias Mertens, Natalia Fischl-Lanzoni, Neil Thompson · February 06, 2026
arxiv correlational medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Matthias Mertens unresolved corpus identity
  2. Natalia Fischl-Lanzoni unresolved corpus identity
  3. Neil Thompson unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Matthias Mertens provider ID
  2. Natalia Fischl-Lanzoni provider ID
  3. Neil Thompson provider ID
At the performance frontier, higher training compute explains 80–90% of model differences, whereas away from the frontier developer-specific methods and shared algorithmic progress substantially reduce required compute and some firms produce far more compute-efficient smaller models, with within-firm efficiency sometimes varying over 40×.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Do leading LLM developers possess a proprietary ``secret sauce'', or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, we estimate scaling-law regressions with release-date and developer fixed effects. We find clear evidence of developer-specific efficiency advantages, but their importance depends on where models lie in the performance distribution. At the frontier, 80-90% of performance differences are explained by higher training compute, implying that scale--not proprietary technology--drives frontier advances. Away from the frontier, however, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Some companies can systematically produce smaller models more efficiently. Strikingly, we also find substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. We also discuss the implications for AI leadership and capability diffusion.

Summary

Main Finding

Scaling (training compute) is the dominant driver of frontier LLM performance, but there is robust evidence of a developer “secret sauce” that materially affects performance away from the frontier. Using 809 publicly released LLMs (Oct 2022–Mar 2025) and MMLU‑Pro scores, the authors decompose performance variation into scaling, shared algorithmic progress (time effects), developer-specific effects (“secret sauce”), and model‑specific residuals. At the frontier, 80–90% of performance differences are attributable to larger training compute; away from the frontier, shared and proprietary algorithmic gains substantially reduce the compute needed to reach fixed capability thresholds.

Key Points

  • Dataset and outcome
    • 809 LLMs released 2022q4–2025q1, benchmarked on MMLU‑Pro (and robustness checks on MATH L5).
  • Empirical approach
    • Regress logit(MMLU‑Pro) on log(training FLOPs), period dummies, and developer dummies.
    • Training compute ci computed as ci = 6 × Ntokens × Nparams (Kaplan et al. approach).
    • Use Shapley variance decomposition to apportion explained variance to scaling, shared algorithmic progress (period), company fixed effects, and residual (model) effects.
  • Main quantitative results (representative figures)
    • Baseline variance shares (all 809 models): scaling ≈ 32%, shared alg. progress ≈ 2.7%, company secret sauce ≈ 14–18% (authors report 14–18% overall), model residuals ≈ 47%.
    • Among major developers only: scaling ≈ 45%, shared ≈ 5%, company ≈ 14%, residuals ≈ 36%.
    • Frontier models: scaling explains ~80–90% of the rise in top‑score benchmarks (example: Llama2 70B → GPT‑4.5). Training compute at the frontier rose ~5,000× across the sample window.
    • Shared algorithmic progress over the window increased effective compute by ≈7.5× (MATH L5 test: ≈11×).
    • Developer (company) heterogeneity: some developers show very large estimated compute efficiency factors (e.g., Microsoft ≈60.5× relative to small dev baseline; DeepSeek ≈2.3× — authors note large uncertainty and heterogeneity).
    • Within‑company/model heterogeneity: model‑specific residuals imply a 41× difference in effective compute between the 90th and 10th percentile of models (MATH L5 shows even larger dispersion, up to 186×).
    • For non‑frontier thresholds: the minimum compute required to reach a 15% normalized MMLU‑Pro score fell ~50× among major developers (≈8,000× if small developers included).
  • Regression/fit
    • Coefficient on log FLOPs: a 10× increase in training compute raises log‑odds of MMLU‑Pro by ~0.79 (SE 0.19). Model R² reported ~0.52 in the main specification.
  • Robustness/auxiliary results
    • Similar qualitative patterns for MATH Level 5, with secret sauce relatively more important for math capabilities and larger model‑specific dispersion.
    • Results robust to alternative developer sets and period specifications (authors report robustness checks).

Data & Methods

  • Outcome variable
    • MMLU‑Pro benchmark scores, transformed via logit: Y = ln(y/(1−y)) to match sigmoidal relationship with log compute.
  • Key regressors
    • log(training compute) — computed from publicly reported parameters and token counts using ci = 6 × Ni × Di.
    • Period fixed effects: three roughly equal time blocks (baseline = 2022q4–2023q3; comparisons for later periods capture shared algorithmic improvements).
    • Developer fixed effects: dummies for main companies (Deepseek, Qwen, Meta, Google, Microsoft, OpenAI, Anthropic, X‑AI, 01‑AI, Nvidia) plus grouped small/medium/large “other” developers.
  • Identification/interpretation
    • Period fixed effects interpreted as aggregate/shared algorithmic/technical efficiency gains (diffuse improvements used by many teams).
    • Developer fixed effects interpreted as persistent developer‑specific efficiency advantages — the “secret sauce” (captures firm‑level reproducible practices, datasets, fine‑tuning pipelines, teacher models, engineering).
    • Model residuals capture model‑specific specializations, fine‑tuning choices, experimentation and measurement noise.
  • Decomposition
    • OLS estimates feed into a Shapley decomposition of explained variance (R²) to assign shares to each factor.
    • Authors also present counterfactual decompositions (e.g., how much of frontier improvement would remain under alternative compute/time/company conditions).
  • Limitations noted by authors
    • Observational design: fixed effects are descriptive and not causal proofs of mechanisms (e.g., a developer fixed effect bundles many possible channels).
    • Measurement: compute is inferred from parameters/tokens with the standard approximation; model performance benchmark choice (MMLU‑Pro) may emphasize some capabilities over others.
    • Sample selection: public/reported models only; unpublished internal models and proprietary compute figures are absent.

Implications for AI Economics

  • Frontier competition depends on compute access
    • Because scale explains most frontier gains, entities that secure superior, expanding compute resources are best positioned to lead in frontier capabilities. Industrial policy and firm strategies that ensure large-scale compute (data‑center capacity, GPUs/accelerators, chilled power, procurement) will be decisive.
  • Secret sauce matters for cost‑efficiency and commercial/product competition
    • Developer‑level efficiency advantages (14–18% of performance variance overall, higher away from frontier) create persistent cost and product differentiation even when scale is limited. Firms with better data, fine‑tuning, pipelines or engineering can deliver equivalent capabilities at lower compute/inference cost — a commercial advantage for productization and margins.
  • Democratization and diffusion below the frontier
    • Shared algorithmic improvements and proprietary efficiency gains have drastically lowered the compute required for given capability thresholds (e.g., 15% MMLU‑Pro). This reduces the barriers for smaller developers and countries to build useful/competent models, accelerating diffusion of many LLM capabilities and lowering inference costs for users.
  • Risk and governance tradeoffs
    • Lower compute requirements for non‑frontier capabilities raise proliferation risks: capable models become easier to train or fine‑tune for misuse. Yet frontier capabilities (most transformative/novel risks) remain concentrated around compute. Governance ought to differentiate policy: controls on large compute supply chains, export controls, or monitoring might target frontier risk, while mitigation of diffusion risks requires monitoring of model releases, datasets, and reuse.
  • Strategic firm behavior and market structure
    • Two competitive levers exist: (1) invest in scale to pursue frontier leadership; (2) invest in algorithmic/data/process improvements to reduce costs and widen product offerings. Both can coexist: firms may pursue frontier scale for headline capabilities while using secret sauce to improve margins and broaden commercial reach.
  • Future of progress if compute growth slows
    • If compute availability growth decelerates, the relative importance of algorithmic and firm‑level efficiency will increase. The observed 7.5× shared efficiency gain over ~2 years demonstrates meaningful scope for continued algorithmic returns; however, current data suggest those gains have so far been insufficient to supplant scale at the frontier.
  • Research and policy priorities
    • Measure and monitor compute concentration and supply chains (because access to compute is central to frontier dominance).
    • Encourage transparency on model training (compute, data provenance) to better assess diffusion and risk.
    • Support public and academic work on reproducible algorithmic efficiency (to democratize capabilities while giving policymakers better signals).
    • Distinguish governance interventions targeted at frontier compute/control versus broader model release practices to address lower‑scale but widespread risks.

Caveats and next steps - The fixed‑effects approach is descriptive; causal inference on which specific practices constitute the “secret sauce” (datasets, fine‑tuning recipes, teacher models, architecture tweaks) requires firm‑level process data or experiments. - The benchmark suite used (MMLU‑Pro, MATH L5) captures particular capability axes; other capabilities may exhibit different scaling/firm effect patterns. - Future work: replicate with private/unreleased models where possible, expand benchmarks, and decompose company effects into measurable components (data quality, prompt‑engineering pipelines, fine‑tuning budgets, etc.).

Assessment

Paper Typecorrelational Evidence Strengthmedium — Large sample (809 released models) and fixed-effect regressions provide substantial correlational evidence that compute explains much of frontier performance, but the analysis is observational: training-compute is not randomly assigned and is potentially endogenous (selection into compute budgets, unobserved data/architecture quality), compute estimates can be noisy, and released-models sample may be biased, all of which limit causal certainty. Methods Rigormedium — Uses well-established scaling-law regressions, developer and time fixed effects, variance decomposition and distributional slices (frontier vs off-frontier) which are appropriate and informative; however, results hinge on functional-form choices, accuracy of compute and benchmark measurements, potential omitted variables (data quality, architecture choices, hyperparameter tuning), and lack of instrumental variables or natural experiments to strengthen causal claims. SampleDataset of 809 publicly released LLMs from 2022–2025 with estimated training compute, benchmark performance metrics across tasks, release dates and developer identities; includes models across many firms and capability levels but restricted to released models and available benchmark scores and compute estimates. Themesinnovation governance IdentificationEstimate scaling-law regressions of benchmark performance on estimated training compute with release-date and developer fixed effects; decompose explained variance to attribute differences to compute versus developer-specific (proprietary) effects and perform distributional analyses (frontier vs non-frontier) and within-firm comparisons. Identification relies on cross-model, observational variation rather than exogenous manipulation. GeneralizabilitySelection bias: only publicly released models included (excludes proprietary/undeclared internal models and withheld failures)., Measurement error: training-compute and some benchmark scores are estimated or inconsistently reported across developers., Benchmark scope: standard benchmarks may not capture all dimensions of model capability or real-world economic performance., Time-limited: data covers 2022–2025 and may not generalize to future architectural or algorithmic breakthroughs., Model types/modalities: focused on LLMs; findings may not extend to other model families or multimodal systems., Firm heterogeneity: internal practices and unobserved inputs (data curation, labeling, evaluation cycles) may limit firm-to-firm comparability.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We use training and benchmark data for 809 models released between 2022 and 2025. Other null_result dataset size / sample of models
Reading fidelity high
Study strength high
n=809
0.5
We estimate scaling-law regressions with release-date and developer fixed effects. Innovation Output null_result relationship between training compute and model performance, controlling for release date and developer
Reading fidelity high
Study strength high
n=809
0.5
There is clear evidence of developer-specific efficiency advantages. Innovation Output positive developer-specific compute efficiency (compute required for given performance)
Reading fidelity high
Study strength medium
n=809
0.3
The importance of developer-specific advantages depends on where models lie in the performance distribution. Innovation Output mixed heterogeneity of developer efficiency across performance quantiles
Reading fidelity high
Study strength medium
n=809
0.3
At the frontier, 80–90% of performance differences are explained by higher training compute, implying that scale — not proprietary technology — drives frontier advances. Innovation Output positive share of performance differences explained by training compute at frontier models
Reading fidelity high
Study strength medium
n=809
80-90% of performance differences explained by higher training compute
0.3
Away from the frontier, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Innovation Output positive compute required to reach fixed capability thresholds outside the frontier
Reading fidelity high
Study strength medium
n=809
0.3
Some companies can systematically produce smaller models more efficiently. Innovation Output positive compute-to-capability efficiency of specific companies
Reading fidelity high
Study strength medium
n=809
0.3
There is substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. Innovation Output mixed within-firm compute-efficiency variation (ratio of compute required for models achieving similar capabilities)
Reading fidelity high
Study strength medium
n=809
more than 40x compute efficiency difference
0.3
We discuss the implications for AI leadership and capability diffusion. Governance And Regulation null_result implications for AI leadership and capability diffusion (qualitative)
Reading fidelity high
Study strength speculative
n=809
0.05

Notes