Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways for European Energy Desks
- EPT-2 delivers state-of-the-art RMSE on four energy-critical variables (10 m wind, 100 m wind, 2 m temperature, surface solar radiation) across the full 0–240 hour horizon.
- The model beats Microsoft Aurora, ECMWF IFS HRES, and other leading NWP and AI systems on European ground-station data from intraday to ten days ahead.
- EPT-2 captures seasonal error patterns such as winter temperature inversions and Alpine orographic effects more accurately, which improves forecasts during high-impact trading periods.
- Better hub-height wind and solar forecasts drive direct P&L gains, with a four-percentage-point accuracy improvement worth roughly €1.5 M–€3 M per year per GW.
- Run a live EPT-2 benchmark against your current provider in under five minutes and see the accuracy gap on your own region.
EPT-2 as the Current Accuracy Leader for Europe
The table below shows EPT-2’s clean sweep across all energy-critical variables and forecast horizons, which establishes it as the current accuracy leader for European energy trading. Lead time is the number of hours between forecast initialization and the valid time of the prediction. Each cell reflects performance aggregated across European ground stations.
| Variable | 0–24 h | 24–72 h | 72–120 h | 120–240 h |
|---|---|---|---|---|
| 10 m wind speed | EPT-2 | EPT-2 | EPT-2 | EPT-2 |
| 100 m wind speed (hub height) | EPT-2 | EPT-2 | EPT-2 | EPT-2 |
| 2 m temperature | EPT-2 | EPT-2 | EPT-2 | EPT-2 |
| Surface solar radiation downwelling | EPT-2 | EPT-2 | EPT-2 | EPT-2 |
EPT-2 outperforms Microsoft Aurora and ECMWF IFS HRES on these energy-relevant variables over the 0–240 h range (arXiv:2507.09703). EPT-2 delivers hourly global weather updates and beats leading AI weather models and traditional numerical baselines on RMSE across all forecast horizons. Microsoft Aurora produces no SSRD output, so the solar radiation comparison pits EPT-2 against other models only.
ECMWF vs GFS vs EPT-2 for European Trading
For European energy applications, the relevant comparison is no longer ECMWF versus GFS alone; it is EPT-2 against the full field. ECMWF HRES has held the deterministic NWP benchmark for forty years, and NOAA GFS is the free global baseline. The EPT-2 technical report shows strong performance relative to both on the listed European variables.
The ECMWF versus GFS framing is also dated, because ECMWF AIFS is now the institution’s primary AI benchmark. ECMWF AIFS, ECMWF’s own AI forecasting system, represents its most direct response to the AI-weather generation. It runs on the Jua platform alongside EPT-2, which allows direct comparison in the same workspace. The report shows EPT-2 outperforming ECMWF IFS HRES across the 0–240 h range on the energy variables.
Microsoft Aurora, the previous state of the art in AI weather before EPT-2, loses to EPT-2 on 10 m wind, 100 m wind, and 2 m temperature across the full 0–240 h range. Aurora produces no SSRD output. DWD ICON, the German Weather Service global and regional model, is the reference for central European grid operators and sits below some leading models in the same evaluation.
Between ECMWF and GFS specifically, ECMWF HRES consistently outperforms NOAA GFS on European variables at medium range, and that pattern holds across multiple independent evaluations. EPT-2 shows strong performance relative to both. For trading desks, model selection is no longer a binary ECMWF-or-GFS decision; EPT-2 is now the accuracy leader on the variables that drive European power and gas P&L.
Seasonal Error Patterns in 2025–2026 StationBench Data
StationBench evaluation across the 2025–2026 period surfaces two recurring seasonal error structures that matter for European energy trading. Both patterns highlight how EPT-2’s training on dense ground-station networks delivers measurable gains during high-impact seasons.
The first pattern is the winter inversion problem. During December–February, stable boundary-layer conditions over central and eastern Europe, particularly in the Po Valley, the Rhine Rift Valley, and the Pannonian Basin, produce persistent temperature inversions. These inversions suppress 10 m wind speeds and decouple surface conditions from the free atmosphere. NWP models with coarser vertical resolution systematically overestimate surface wind speeds under inversion, which inflates RMSE on 10 m wind relative to the annual mean. EPT-2’s training on more than 10,000 ground stations, including dense European surface networks, captures the observational signature of these inversions directly. That effect contributes to its winter-season RMSE advantage over ECMWF HRES and NOAA GFS on near-surface wind.
The second pattern is Alpine orographic performance. The Alpine arc, from the Swiss Plateau through the Austrian Vorarlberg to the Slovenian Karawanks, generates föhn events, gap flows, and lee-wave structures that models at 9 km resolution, such as ECMWF HRES, or coarser grids under-resolve. These structures drive sharp wind ramps and temperature spikes that carry high value for Swiss and Austrian hydro and wind operators. EPT-2’s native forecast capability down to 5 km resolution over Europe through EPT-2 HRRR narrows the orographic RMSE gap that persists in continental-scale models at standard resolution. Customers with Alpine assets, including Axpo and Statkraft, run EPT-2 HRRR alongside the global EPT-2 deterministic run for this reason.
Energy P&L Impact of Lower Hub-Height Wind RMSE
Hub-height wind, typically measured at 100 m, is the single variable with the largest direct P&L impact for European wind operators. A one-percentage-point improvement in wind forecast accuracy at day-ahead reduces imbalance exposure and hedging cost. The market-sizing economics are concrete: a 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 M per year under typical European hedging and imbalance penalty structures. A 1 GW solar portfolio at the same accuracy gain saves approximately €3 M per year, which reflects the higher intraday price volatility of solar generation.
EPT-2’s RMSE lead on 100 m wind across the full 0–240 h range fits directly into this economics frame. The advantage is most pronounced at 72–120 h, the lead time that governs day-ahead and week-ahead position-taking, because medium-range NWP error compounds most severely there. EPT-2 avoids this compounding through its native any-Δt architecture. Unlike Aurora and most peers, which roll forward in fixed 6-hour increments and accumulate error at each step, EPT-2 is trained to predict at arbitrary lead times directly, which removes step-wise error propagation.
For multi-GW portfolios, such as those operated by Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec, these economics scale linearly. A 3 GW wind book that moves from ECMWF HRES to EPT-2 as its primary forecast signal, and captures four percentage points of accuracy gain, targets roughly €4.5 M in annual savings before any additional value from intraday alert capture.
Quantify the P&L impact by benchmarking EPT-2 against your current feed on your own regions and variables in under five minutes.
How StationBench Evaluates Weather Models
StationBench is an open-source benchmarking protocol that evaluates NWP and AI weather model outputs directly against real ground-station observations, with no post-processing, bias correction, or station fine-tuning applied to any model. The evaluation covers ground stations across Europe, drawn from SYNOP surface networks, METAR aviation observations, and proprietary station feeds. All models are initialized from the same operational analysis fields and evaluated at matching valid times.
Key metric definitions for this report:
- RMSE (root mean square error): the square root of the mean of squared differences between forecast and observed values. Lower values are better and the metric is sensitive to large errors.
- CRPS (continuous ranked probability score): a proper scoring rule for probabilistic forecasts that measures both accuracy and calibration. Lower values are better.
- NWP (numerical weather prediction): the class of physics-based forecast models that solve atmospheric differential equations on a three-dimensional grid. ECMWF HRES, NOAA GFS, and DWD ICON are NWP models.
- Ensemble: a set of model runs initialized with perturbed conditions that produces a probability distribution over forecast outcomes. EPT-2e is Jua’s ensemble variant; ECMWF ENS is the 50-member NWP reference.
- Lead time: the number of hours between forecast initialization and the valid time of the predicted value.
EPT-2e, the ensemble variant of EPT-2, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time (arXiv:2507.09703). The full StationBench methodology is documented on the Jua StationBench methodology page.
Conclusion: EPT-2 as a Defensible Benchmark for Model Selection
The EPT-2 technical report confirms EPT-2 as the accuracy leader on every energy-critical European variable across the full 0–240 h forecast horizon. For trading desks that manage multi-GW portfolios, the four-percentage-point accuracy gain translates into the €1.5 M–€3 M per GW annual savings cited earlier, which makes model selection an economic decision rather than a branding choice.
Meteorologists, quant developers, and trading teams now have a single citable benchmark to defend model-selection decisions: arXiv:2507.09703. The same benchmark is reproducible live on the Jua platform in under five minutes on any region and variable that matters to your book.
Verify these results yourself by running EPT-2 head-to-head against ECMWF HRES, ECMWF AIFS, NOAA GFS, DWD ICON, and Microsoft Aurora on your own regions and variables.
Frequently Asked Questions
What does RMSE measure in weather forecasting, and why does it matter for energy trading?
RMSE, or root mean square error, measures the average magnitude of forecast errors by taking the square root of the mean of squared differences between predicted and observed values. Because it squares errors before averaging, RMSE penalizes large deviations more heavily than small ones. That property makes it suitable for energy applications where a single large forecast miss, such as an unpredicted wind ramp or solar dip, can move P&L more than many small errors combined. For a 1 GW wind portfolio, a four-percentage-point improvement in forecast accuracy matches the €1.5 M annual savings figure cited earlier. For solar, the figure is €3 M per GW because intraday price volatility is higher. RMSE evaluated against real ground stations, as StationBench does, offers the most defensible basis for model selection because it reflects performance on the observations that drive grid and market outcomes, not on smoothed reanalysis fields.
How is EPT-2 different from ECMWF HRES, and does Jua for Energy replace an ECMWF subscription?
EPT-2 is the deterministic flagship of the EPT family, a general spatiotemporal foundation model trained on more than 5 petabytes of observational physics data from over 120 sources. It outperforms Microsoft Aurora and ECMWF IFS HRES on the four energy-critical variables across the full 0–240 hour range, as documented in the EPT-2 technical report. The architectural difference is significant. EPT-2 is trained to predict at arbitrary lead times, a native any-Δt design, rather than rolling forward in fixed 6-hour increments. That design removes the error compounding that affects most NWP and AI peer models at medium range.
Jua for Energy does not replace an ECMWF subscription. Serious customers keep their ECMWF feed and run EPT-2 alongside it. ECMWF AIFS, ECMWF’s own AI model, runs natively on the Jua platform in the same workspace as EPT-2. Jua for Energy instead replaces the plumbing around the incumbent feed, including the in-house grib pipeline, manual benchmarking, morning-briefing routine, and spreadsheet stitching that often consumes the 7–9 a.m. window before the market opens.
What is StationBench, and why is it more reliable than vendor-provided accuracy graphics?
StationBench is an open-source benchmarking protocol that evaluates weather model outputs directly against real ground-station observations with no post-processing, bias correction, or station fine-tuning applied to any model. The key distinction from vendor-provided accuracy graphics is independence and reproducibility. Every model in the comparison is initialized from the same operational analysis fields, evaluated at matching valid times, and scored against the same observation set. No model receives preferential treatment through post-processing.
The protocol is publicly documented, and the results appear in peer-reviewed form at arXiv:2507.09703. On the Jua platform, the same benchmark is reproducible live in under five minutes on any region and variable a customer selects. That live benchmark moment is often when meteorologists who were sceptical of vendor accuracy claims become internal champions, because they run the numbers themselves instead of relying on pre-selected graphics.
What is the difference between EPT-2 and EPT-2e, and when should a trading desk use each?
EPT-2 is the deterministic flagship, a single best-estimate forecast initialized four times per day, running out to 20 days at up to 5 km resolution over Europe. It is the model that records state-of-the-art performance on all four energy-critical European variables according to the EPT-2 technical report. EPT-2e is the ensemble variant. It produces a distribution of forecast outcomes instead of a single trajectory, which enables probabilistic risk quantification. EPT-2e beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time.
For a trading desk, the split is straightforward. EPT-2 is the primary signal for point-estimate positioning, including day-ahead generation forecasts, price-curve inputs, and dispatch decisions. EPT-2e is the signal for probabilistic risk management, including tail-event sizing, option pricing, imbalance reserve calculations, and any position where the width of the forecast distribution matters as much as the central estimate. Quant funds typically pipe both through the Jua Python SDK and combine them in their own systematic models.
How quickly can a meteorologist or quant team verify EPT-2’s accuracy on their own region?
The live benchmarking surface on the Jua platform returns a head-to-head accuracy comparison in under five minutes. A user selects a region, such as any European country, sub-national zone, or custom bounding box, a variable, and a time window, then adds EPT-2 alongside their current provider. The platform runs the comparison across all selected models and returns RMSE and CRPS results against the StationBench ground-station network.
No data engineering, pipeline setup, or vendor coordination is required. Backtests against years of historical forecasts run in roughly five minutes through Athena, Jua’s AI agent, or directly through the Python SDK for teams that prefer programmatic access. For most Jua for Energy customers, the benchmark is the deal trigger, because the conversation shifts from scepticism about accuracy claims to implementation timing once the numbers appear on the customer’s own region and variable.
