Written by: Olivier Lam, Physical AI Team, Jua.ai AG
What You Will Learn From This Benchmark Guide
- Renewable energy forecasting accuracy directly drives European power prices, with forecast errors creating intraday volatility and substantial balancing penalties for BRPs.
- Spatiotemporal foundation models like EPT-2 beat legacy NWP benchmarks such as ECMWF HRES across wind, temperature, and solar radiation variables at every lead time from 0–240 hours.
- EPT-2 delivers forecasts at up to 1 km resolution with update frequencies up to 24 times per day at a fraction of the compute cost of traditional NWP simulations.
- Jua for Energy combines EPT-2 physics forecasts with Athena agents to deliver actionable briefings, benchmarks, and backtests that plug into existing trading pipelines.
- Benchmark EPT-2 against your current provider and unlock higher-accuracy renewable forecasts for European energy trading.
Why Forecast Accuracy Moves European Power Prices
Weather is the single largest driver of short-term electricity prices in Europe. Forecast errors in wind and solar generation create intraday price volatility when actual output deviates from predicted levels, with revisions in forecasts, not absolute forecast levels, driving the swings. Balance Responsible Parties (BRPs) must rebalance portfolios at the last minute in the intraday market to avoid large balancing penalties when renewable forecasts are inaccurate.
The financial stakes are quantifiable. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 million per year under typical hedging and penalty structures. A 1 GW solar portfolio at the same accuracy gain saves approximately €3 million per year. Operators running multi-GW portfolios scale these economics linearly. The EU reached 406 GW of solar capacity by the end of 2025 and targets at least 600 GW by 2030, which accelerates the cannibalisation effect, widens the distribution of hourly prices, and raises the cost of forecast error. Managing this risk requires clear alignment between forecast horizons and trading decisions.
Quantify the savings for your portfolio. See EPT-2 benchmarked against your current provider.
How Short- and Long-Term Horizons Support European Trading
Lead time, the interval between when a forecast is issued and the moment it describes, determines which market decision it supports. Short-term forecasting covers 0–48 hours and maps directly to intraday trading. As of October 2025, 15-minute intervals are the standard Market Time Unit across European power markets, which enables continuous position refinement as weather forecasts improve. A forecast made 24 hours in advance is rarely accurate enough to avoid residual imbalances due to rapid solar production ramping within a single quarter-hour window. High-frequency short-range updates therefore become operationally critical.
Long-term forecasting spans day-ahead through 20 days and supports TSO capacity planning, portfolio hedging, and multi-day position management. Day-ahead prices serve as the primary reference through centralised auctions, yet forecast errors, renewable intermittency, and demand uncertainty leave substantial risk close to delivery. Transmission system operators rely on accurate medium-range outlooks to schedule balancing reserves. Traders use them to position around expected generation deficits or surpluses before the day-ahead gate closes.
Machine Learning Models Now Power European Solar Irradiance Forecasts
Legacy numerical weather prediction (NWP), the method of decomposing the atmosphere into three-dimensional grid cells and solving differential equations inside each one, has produced reliable solar irradiance forecasts for decades. The constraint is compute: a single NWP simulation consumes approximately 8,400 kWh and costs €1,000–€20,000 to run, which caps update frequency at two to four runs per day.
Spatiotemporal foundation models learn the governing physics of complex systems directly from observational data, then run inference on a single GPU in minutes at approximately 0.25 kWh and $0.20–$15 per simulation. EPT-2, Jua’s general physics foundation model fine-tuned for atmospheric prediction, produces forecasts at native any-Δt, meaning it is trained to predict at arbitrary time steps rather than rolling forward in fixed 6-hour increments. This architectural choice matters because most AI peers, including Microsoft Aurora, roll forward in 6-hour steps, which compounds error across the forecast horizon. EPT-2’s non-rolling approach preserves accuracy at every lead time, and it does so at up to ~5 km native resolution over Europe through the EPT2-HRRR variant, covering surface solar radiation (SSRD) alongside wind at 11 height levels from 10 m to 200 m.
A 2026 study published in Renewable Energy reports that Bayesian-optimised multilayer perceptrons can achieve competitive day-ahead performance for wind and solar power in Germany when trained on reanalysis data and evaluated against NWP outputs including ERA5, ICON, IFS, and GFS. These results establish the baseline that physics foundation models are now measured against.
How Forecast Errors Feed Into EU Electricity Prices
Intraday electricity prices in Europe are driven by micro-orderbook dynamics and rapidly updating forecasts, which renders them more sensitive to forecast revisions than day-ahead prices. Repricing hotspots occur 1–3 hours before market closing, often the most frantic period, while 3 p.m. and evening slots offer the highest liquidity with shorter 15-minute updates.
Balancing prices in European markets are tightly linked to system imbalance and TSO reserve activation rules, producing the highest volatility among day-ahead, intraday, and balancing market types. During wind droughts, renewable supply deficits require procurement of more costly power, which drives price increases. Excess renewable generation with insufficient demand forces energy to be sold at lower prices, which causes price drops. The asymmetry between these two outcomes makes probabilistic forecasting, which quantifies the range of likely outcomes rather than committing to a single point estimate, operationally essential for BRPs managing exposure across a renewables portfolio.
EPT-2 vs ECMWF HRES: European Wind and Solar Benchmarks
ECMWF HRES (High RESolution forecast) is the deterministic NWP benchmark that has led global atmospheric prediction for over forty years. EPT-2, documented in arXiv:2507.09703 and building on the methodology established in arXiv:2410.15076, outperforms ECMWF HRES on every lead time across the full 0–240 hour range on four variables that directly drive energy P&L: 10 m wind speed, 100 m wind speed, 2 m temperature, and surface solar radiation (SSRD). These benchmarks are evaluated against more than 10,000 real ground stations using the open-source StationBench methodology, with no post-processing or station fine-tuning applied.
EPT-2e, the ensemble variant of EPT-2, beats the 50-member ECMWF ENS mean on both RMSE (root mean square error, the standard deviation of forecast residuals) and CRPS (Continuous Ranked Probability Score, a proper scoring rule that evaluates the full forecast distribution against observations) at virtually every lead time. EPT-2e updates 4 times per day. On the Jua platform, Jua for Energy’s customer-facing product surface, forecasts are available at up to 1 km resolution.
EPT-2 also outperforms Microsoft Aurora on 10 m wind, 100 m wind, and 2 m temperature across the full 0–240 hour range. Aurora produces no SSRD output, so EPT-2 wins that variable by default. Inference on EPT-2 runs approximately 25% faster than Aurora.
Run these benchmarks on your own region and variables. Request access to the Jua platform.
Why Foundation Models Beat Legacy NWP for 2030 Grid Targets
Europe’s 2030 renewable targets require grid operators and traders to manage a generation mix that is increasingly weather-dependent. The table below compares EPT-2 and EPT-2e against ECMWF HRES/ENS, Microsoft Aurora, and Google DeepMind GraphCast across four operationally critical dimensions. All benchmark figures are sourced from arXiv:2507.09703 and arXiv:2410.15076. Operational specifications come from Jua’s published product documentation.
| Model | Deterministic accuracy (0–240 h, 10 m wind / 100 m wind / 2 m temp / SSRD vs HRES) | Ensemble skill (RMSE & CRPS vs ECMWF ENS mean) | Update frequency |
|---|---|---|---|
| EPT-2 / EPT-2e (arXiv:2507.09703) | Outperforms ECMWF HRES on all four variables across full 0–240 h range, outperforms Aurora on 10 m wind, 100 m wind, 2 m temp across full range, wins SSRD by default (Aurora has no SSRD output) | EPT-2e beats 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time (arXiv:2507.09703) | 4×/day (EPT-2e), up to 24×/day (EPT-2 RR) |
| ECMWF HRES / ENS | 40-year NWP benchmark, definitive reference for traders repricing risk around heating demand, renewable output, and system tightness | ENS: 50-member gold standard for probabilistic NWP | 2–4×/day |
| Microsoft Aurora | Loses to EPT-2 on 10 m wind, 100 m wind across full 0–240 h range, loses on 2 m temp, no SSRD output (arXiv:2507.09703) | No productised ensemble equivalent | Typically 4×/day research cadence, no productised operational schedule |
| Google DeepMind GraphCast | EPT-1.5 outperforms GraphCast on European wind and temperature (arXiv:2410.15076) | No productised ensemble equivalent | Typically 4×/day research cadence, no productised operational schedule |
The infrastructure cost asymmetry compounds the operational advantage. EPT-2 was trained on 8 × H100 GPUs over 10 days, while Microsoft Aurora required 32 × A100 GPUs over 18 days. At inference, a single EPT-2 run costs approximately $0.20–$15 at ~0.25 kWh on a single GPU, versus €1,000–€20,000 at ~8,400 kWh for a traditional NWP simulation on HPC. That four-orders-of-magnitude cost advantage, the ~€10,000 NWP run versus the ~$10 EPT-2 run, is what makes 24 daily refreshes economically viable without an HPC cluster.
How Jua for Energy Turns Physics Into Trading Decisions
Jua is a foundation model and agent company. The relationship to Jua for Energy mirrors the relationship Anthropic has to Claude Code: a horizontal AI platform, EPT for physics and Athena for agency, with a flagship vertical product applied to energy trading. EPT and Athena are domain-agnostic by architecture. The atmosphere is the first physical system EPT has been fine-tuned for, and energy trading is the first market Athena has been instrumented for.
Inside Jua for Energy, Athena turns raw physics predictions from EPT-2 into actionable briefings by reading market context and modelling participant behaviour. A typical natural-language query resolves in approximately 90 seconds. A backtest runs in approximately 5 minutes. Day-Ahead and Intraday briefings auto-refresh on every new model run, covering model consensus across 25+ models, model delta since the previous run, convergence tracking, and price implications. This workflow replaces the manual 7–9 a.m. routine of downloading grib files, stitching together spreadsheets, and juggling terminal screens.
Power forecasts cover solar, wind onshore, wind offshore, total wind, total renewables, load, and residual load across five countries: Germany, Great Britain, France, the Netherlands, and Belgium. Actual generation refreshes every 15 minutes. The Fundamental Model runs out to 20 days. Divergence alerts fire the moment two models disagree on a key variable. Correction alerts fire the moment a model revises its own output, surfacing trade windows before the market re-prices.
For quant developers and engineering teams, pip install jua installs the Python SDK. The REST API exposes 25+ models, 10 proprietary AI from the EPT family plus 15 third-party NWP and AI models including ECMWF HRES, ENS, AIFS, NOAA GFS, DWD ICON, Aurora, and GraphCast, through a single schema with Apache Arrow support for large payloads. Hindcast data is available across multiple Jua and third-party models for backtesting. Jua serves major utilities across four continents, including some of Europe’s largest energy companies, as well as commodity traders and hedge funds, among them Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec.
Frequently Asked Questions
How do forecast errors affect European intraday prices?
Forecast errors in wind and solar generation create intraday price volatility when actual output deviates from predicted levels. The mechanism is BRP rebalancing: when a renewable forecast is inaccurate, Balance Responsible Parties must buy or sell power at the last minute in the intraday market to avoid balancing penalties. These last-minute transactions move prices upward when generation falls short of forecast and additional power must be procured, and downward when generation exceeds forecast and surplus power must be offloaded. Since October 2025, 15-minute Market Time Units are the European standard, so forecast errors now propagate into prices at a finer granularity than ever before. Probabilistic forecasts, which quantify the range of likely outcomes rather than committing to a single point estimate, give traders a clearer picture of exposure and allow more precise hedging of residual imbalance risk.
What resolution and update frequency does EPT-2 provide for Europe?
EPT2-HRRR natively forecasts at up to ~5 km resolution over Europe. On the Jua platform, the customer-facing product surface of Jua for Energy, forecasts are available at up to 1 km resolution. EPT-2e, the ensemble variant, updates 4 times per day. EPT-2 RR, the rapid-refresh variant, updates up to 24 times per day. Actual-generation power forecasts refresh every 15 minutes. This update cadence contrasts with traditional NWP, which is economically constrained to 2–4 runs per day by the cost of HPC infrastructure. EPT-2 inference runs on a single GPU in minutes at approximately 0.25 kWh and $0.20–$15 per simulation, which makes high-frequency refreshes viable without an HPC cluster.
How does EPT-2 compare with ECMWF HRES on wind and solar variables?
EPT-2 beats ECMWF HRES across all lead times on the four variables that matter most for energy trading: wind at 10 m and 100 m, temperature at 2 m, and surface solar radiation. Full benchmark results are documented in the EPT-2 technical report at arXiv:2507.09703, evaluated against more than 10,000 real ground stations using the open-source StationBench methodology with no post-processing or station fine-tuning. EPT-2e, the ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. Jua for Energy does not replace ECMWF. Serious customers keep their ECMWF subscription and run Jua for Energy alongside it. What the platform displaces is the plumbing around the incumbent feed: the in-house grib pipeline, the manual benchmarking, and the morning-briefing assembly routine.
Can Jua forecasts integrate with existing trading pipelines?
Yes. Jua for Energy exposes a REST API with Apache Arrow payload format and a Python SDK installable via pip install jua from PyPI. The API provides forecast access, hindcast and backtesting data, and weather-parameter standardisation across all 25+ models on the platform under a unified schema. ENTSO-E grid data, including actual generation, capacity, and PSR classifications across European power markets, integrates directly. Quant developers pipe Jua forecasts into their own systematic models. Utilities and trading houses pipe them into existing dispatch, risk, and trading tools. The integration that takes a quarter to build from raw AI-weather research outputs stands up in days on the Jua platform. Documentation is available at docs.jua.ai and the developer dashboard at developer.jua.ai.
Conclusion: Turning Forecast Accuracy Into P&L
Renewable energy forecasting accuracy in Europe now sits at the centre of P&L. With 15-minute Market Time Units standard, BRP rebalancing penalties accumulating at finer granularity, and solar and wind capacity expanding toward 2030 targets, the cost of a stale or inaccurate forecast compounds faster than legacy NWP infrastructure can refresh.
EPT-2 outperforms ECMWF HRES on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation across the full 0–240 hour lead-time range, as documented in arXiv:2507.09703. EPT-2e beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time. EPT-2 RR updates up to 24 times per day at a fraction of the compute cost of a single NWP run. Athena resolves natural-language queries into briefings, benchmarks, and backtests in approximately 90 seconds. Together, EPT-2 and Athena form the foundation of Jua for Energy, the first applied product from Jua, a foundation model and agent company whose architecture is domain-agnostic and whose roadmap extends well beyond the atmosphere.
The economics are quantifiable: the €1.5 million annual savings for a 1 GW wind portfolio at four percentage points of accuracy gain scales linearly across multi-GW operations. The benchmark is live, transparent, and runs in seconds on your own region and variable.
See EPT-2 head-to-head against your current forecast provider. Schedule your benchmark.
