Pretrained foundation models rarely displace specialized electricity-price forecasters: only TabPFN variants outscore market-specific benchmarks statistically across Germany, Poland and Spain (2021–25). However, statistical superiority does not guarantee higher profits from battery arbitrage—risk-averse strategies favor a distributional deep neural-net benchmark.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland, and Spain over 2021-2025. Their performance is assessed in terms of point and probabilistic forecasting accuracy, as well as economic value in battery energy storage arbitrage. Only the TabPFN models consistently and significantly outperform the benchmarks across all three markets and all statistical measures. However, this statistical dominance does not translate directly into economic dominance: TabPFN performs best under unlimited bids and riskier quantile-based strategies, whereas the Distributional Deep Neural Network benchmark is more profitable when risk tolerance is lower. Thus, foundation models cannot universally replace market-specific models, and their value depends on both model architecture and the decision problem.
Summary
Main Finding
Pretrained foundation models evaluated in zero-shot mode can be competitive with market-specific electricity price forecasting (EPF) models, but they do not universally replace them. Across three European day-ahead markets (Germany, Poland, Spain) over 2021–2025, only the TabPFN family consistently and significantly outperformed two state-of-the-art EPF benchmarks on statistical measures (MAE, RMSE, CRPS). However, statistical superiority did not uniformly translate into greater economic value in Battery Energy Storage System (BESS) arbitrage: TabPFN yields the highest profits under riskier/unlimited bidding strategies, while the Distributional Deep Neural Network (DDNN) benchmark is more profitable under lower risk tolerance.
Key Points
- Scope: Nine FM variants from five families (Chronos-2 variants, Moirai-2, TimesFM-2.5, Mitra, TabPFN variants) compared in zero-shot mode to two EPF benchmarks (LEAR with conformal prediction; DDNN-JSU).
- Markets & period: Day-ahead markets for Germany (DE), Poland (PL), and Spain (ES); five-year out-of-sample test 2021–2025.
- Inputs: Models used market-relevant exogenous covariates where supported — day-ahead load and forecasts of onshore/offshore wind and solar — plus common fundamental variables (EUA, TTF gas, Brent oil, API2 coal futures). Time series were DST-adjusted and preprocessed for missing values.
- Evaluation metrics:
- Point forecasts: MAE, RMSE.
- Probabilistic forecasts: CRPS (and related calibration checks).
- Economic evaluation: BESS arbitrage profit under two decision regimes — quantile-based bidding (risk-sensitive) and unlimited-bid strategies.
- Main empirical results:
- Statistically (MAE / RMSE / CRPS): TabPFN models were the only FM family that consistently and significantly outperformed both EPF benchmarks across all three markets.
- Economically (BESS arbitrage): TabPFN dominated for riskier/unconstrained strategies; DDNN-JSU delivered higher profits under lower-risk (more conservative quantile) strategies. Thus, statistical gains do not map one-to-one to economic gains.
- Mechanistic insight: FMs that natively support exogenous covariates performed better; purely univariate pretrained models can be disadvantaged in EPF where covariates (load, renewables forecasts) are crucial.
- Broader take: Foundation models are a promising component in EPF toolkits but are not a universal plug-and-play substitute for market-specific models or carefully designed EPF pipelines.
Data & Methods
- Data sources: ENTSO-E (day-ahead prices; day-ahead load, onshore/offshore wind, solar generation forecasts) and Investing.com (futures: EUA, TTF gas, Brent, API2 coal). Hourly aggregation from 2017-01-06 to 2025-12-31; test window 2021-01-01 to 2025-12-31.
- Markets chosen to cover structural heterogeneity: high RES shares (Germany, Spain), coal-heavy mix (Poland), presence of nuclear (Spain).
- Foundation models tested (zero-shot):
- Time-series FMs: Chronos-2, Chronos-2-synth, Chronos-2-small, Moirai-2, TimesFM-2.5.
- Tabular / prior-fitted FMs: TabPFN-2, TabPFN-3, TabPFN-TS-3, Mitra (with conformal prediction).
- Benchmarks:
- LEAR (Linear Exponential Additive Regression) with conformal prediction: EPF-specific, efficient baseline for both point and probabilistic prediction.
- DDNN-JSU (Distributional Deep Neural Network with Johnson-SU output): strong probabilistic EPF benchmark producing full predictive distributions.
- Forecasting setup:
- Zero-shot inference only (no fine-tuning), with models that can accept covariates provided covariate support exists.
- Point and probabilistic forecasts produced where models support probabilistic outputs; where necessary, conformal postprocessing used to obtain calibrated intervals/quantiles.
- Economic evaluation:
- Simulated arbitrage with a BESS across the day-ahead market using forecast-driven bid rules:
- Quantile-based strategies: bids based on specific forecast quantiles (risk-sensitive).
- Unlimited-bid strategy: no bid volume constraints; exploits forecast extremes.
- Profit computed across the test period and compared between models.
- Statistical testing: Conditional predictive ability tests and significance assessments comparing FM variants to benchmarks across markets and metrics.
Implications for AI Economics
- Forecast accuracy vs. economic value: Improvements in point/probabilistic accuracy (e.g., CRPS) do not automatically imply greater economic returns. The downstream decision rule (here, BESS arbitrage policy and risk tolerance) mediates how forecast improvements translate to profit. Economic evaluations must be an integral part of model comparison in AI-for-economics contexts.
- Model architecture and task alignment matter: TabPFN’s tabular / prior-fitted approach yielded consistent statistical gains, underscoring that pretraining plus an architecture aligned to the prediction task and covariate usage can outperform large generic time-series FMs. Univariate-trained FMs may underutilize essential covariates and lose ground.
- Risk preferences and strategy design: Which model is most valuable depends on the agent’s risk tolerance. Decision-makers should select models conditioned on the operational policy (e.g., aggressive arbitrage vs. conservative bidding), not solely on statistical scores.
- Role of zero-shot FMs: Zero-shot FMs can be useful when quick deployment or transfer across markets is required, particularly if the FM supports exogenous variables. But they are not yet a universal replacement for market-specific, economically-aware models; fine-tuning, hybrid pipelines, or combining FMs with domain-specific models may be necessary.
- Practical deployment considerations: Evaluate FMs on long periods and across structurally different markets; include realistic covariates and proper preprocessing. Also weigh computational costs of large FMs and the potential benefits of smaller/targeted architectures (e.g., TabPFN) that may offer better cost–benefit trade-offs in EPF tasks.
- Directions for future research and practice:
- Fine-tuning FMs on market-specific data and incorporating spatial/topological information (transmission constraints) could further improve economic performance.
- Hybrid systems (FM-generated covariates feeding simpler downstream models) and temporal reconciliation across product types are promising.
- Routine inclusion of economic evaluation—varying trading rules, storage constraints, and risk profiles—should become standard when assessing forecasting models intended for operational use in energy markets.
Assessment
Claims (5)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Only the TabPFN foundation-model family consistently and significantly outperforms the two market-specific electricity-price-forecasting benchmarks across Germany, Poland, and Spain and across all reported statistical measures. Output Quality | positive | Point and probabilistic electricity-price forecast accuracy |
Reading fidelity
high
Study strength
medium
|
n=3
|
| Foundation models with native support for exogenous variables perform competitively against market-specific electricity-price-forecasting benchmarks, but only the TabPFN family achieves superior statistical accuracy in every market considered. Output Quality | mixed | Statistical accuracy of point and probabilistic electricity-price forecasts |
Reading fidelity
high
Study strength
medium
|
n=3
|
| TabPFN produces the highest battery-arbitrage profitability under unlimited-bid strategies and under riskier quantile-based trading strategies. Firm Revenue | positive | Profit from battery energy storage system electricity-price arbitrage |
Reading fidelity
high
Study strength
medium
|
n=3
|
| The Distributional Deep Neural Network benchmark is more profitable than TabPFN when battery-arbitrage decisions use lower risk tolerance. Firm Revenue | positive | Battery-arbitrage profitability under lower-risk trading strategies |
Reading fidelity
high
Study strength
medium
|
n=3
|
| Foundation models cannot universally replace market-specific electricity-price-forecasting models because their value depends on both model architecture and the downstream decision problem. Task Allocation | mixed | Combined forecasting performance and economic value in battery-storage arbitrage |
Reading fidelity
high
Study strength
medium
|
n=3
|