Written by: Olivier Lam, Physical AI Team, Jua.ai AG
What Energy Desks Can Take From The 2026 Benchmarks
- ECMWF HRES remains the benchmark for European medium-range forecasts and beats NOAA GFS on temperature, wind, and precipitation from 3 to 7 days.
- Jua’s EPT-2 family now surpasses HRES on every lead time and key energy variable, verified against more than 10,000 ground stations using open-source StationBench methodology.
- HRES benefits from 9 km resolution and 4D-Var data assimilation, but Jua’s foundation-model approach now exceeds these structural advantages.
- Leading AI models such as Microsoft Aurora, DeepMind GraphCast, and ECMWF AIFS trail both HRES and EPT-2 in station-level RMSE and CRPS across wind, temperature, and solar radiation.
- Energy traders can run head-to-head benchmarks of HRES, GFS, and EPT-2 on their own regions in minutes. See the impact on your portfolio.
Why The European ECMWF Model Beats The US GFS Over Europe
ECMWF HRES consistently outperforms NOAA GFS over European domains at medium-range lead times. The performance gap is most pronounced from day 4 onward, where HRES maintains higher anomaly correlation and lower RMSE on 2 m temperature and 10 m wind across Central and Northern Europe. Two structural advantages explain this gap.
Resolution. HRES operates at approximately 9 km horizontal grid spacing. GFS runs at 13 km globally. Over complex European terrain, including the Alps, Pyrenees, and Scandinavian fjords, the finer HRES grid resolves orographic forcing that GFS smooths over. This directly improves downstream wind and precipitation skill.
Data assimilation. HRES uses 4D-Var (four-dimensional variational) data assimilation, which optimises the model’s initial state across a time window rather than a single snapshot. GFS uses a hybrid ensemble-variational scheme. Over Europe, where the observational network is dense with SYNOP stations, radiosondes, commercial aircraft, and Meteosat geostationary data, 4D-Var extracts more information from each observation. The result is a more accurate analysis that propagates forward into the forecast.
ECMWF’s two-week outlook is the definitive reference point for traders repricing risk around heating demand, renewable output, and system tightness. GFS has not challenged this status in European markets. For energy desks, HRES is the model to beat, not GFS.
GFS still adds value as a second opinion. Model divergence between HRES and GFS at day 5–7 is itself a tradeable signal. When the two models disagree on a temperature or wind outcome, the spread quantifies forecast uncertainty before it shows up in the price.
How ECMWF HRES Performs On Energy-Critical Variables
ECMWF HRES accuracy depends on variable, lead time, and European sub-region. The points below focus on the variables that drive energy P&L.
2 m temperature. HRES is highly skilful at short range (days 1–3) and retains useful skill through day 7 over most of Europe. Skill degrades faster over the Iberian Peninsula and the Mediterranean. Mesoscale sea-breeze and orographic interactions in these regions are harder to resolve at 9 km.
10 m and 100 m wind. Wind skill at 10 m is strong through day 5 over open terrain and offshore. At 100 m, the hub height relevant to modern wind turbines, HRES does not natively output this variable in all configurations and often requires post-processing. Coastal and complex-terrain sites show higher RMSE because surface roughness transitions remain unresolved at 9 km.
Precipitation. HRES outperforms GFS on stratiform precipitation over Northern and Central Europe. Known limitations include convective precipitation in summer, particularly over the Alps and Balkans, where the 9 km grid cannot resolve convective initiation explicitly. Fog and stratus over the North Sea and English Channel also pose challenges, as boundary-layer parameterisation introduces systematic biases in low-cloud onset and dissipation timing.
Surface solar radiation (SSRD). HRES SSRD skill closely tracks cloud-cover accuracy. Fog and stratus biases propagate directly into SSRD errors. Solar generation forecasts therefore inherit the same boundary-layer limitations.
The ENS role. ECMWF also operates ENS (Ensemble Numerical System, the 50-member probabilistic forecast), which quantifies forecast uncertainty by running 50 perturbed versions of the model. CRPS (Continuous Ranked Probability Score, where lower is better) on ENS is the gold standard for probabilistic skill. For energy trading, ENS spread at day 5–7 is a primary input to risk-management decisions around renewable generation and gas demand.
RMSE (Root Mean Square Error, where lower is better) and CRPS together characterise both deterministic and probabilistic skill. Lead time refers to the number of hours between forecast initialisation and the valid time of the prediction.
What The GFS vs ECMWF Gap Means For Traders
Over Europe, ECMWF HRES is more accurate than NOAA GFS at every lead time from day 1 through day 10 on 2 m temperature, 10 m wind, and precipitation. The gap widens with lead time. By day 7, HRES anomaly correlation on 500 hPa geopotential height over Europe is measurably higher than GFS. This difference translates into better temperature and wind forecasts at the surface.
The resolution advantage at 9 km versus 13 km and the 4D-Var data assimilation advantage are the two primary structural causes. Ensemble design adds a third factor. ECMWF ENS uses singular-vector and stochastic perturbations calibrated over decades. The GFS ensemble uses a different perturbation scheme that produces wider, less calibrated spread over European domains.
For energy trading applications, HRES is the correct primary deterministic reference for European power, gas, and renewables markets. GFS is most useful as a divergence signal. When GFS and HRES disagree materially at day 5–7, that disagreement provides information about forecast uncertainty rather than a reason to prefer GFS.
Compare HRES and GFS directly against EPT-2 on your highest-stakes European region.
Live Multi-Model Verification: HRES, GFS, Aurora, GraphCast, AIFS, And Jua EPT-2
The verification results show a clear pattern: EPT-2 is the only model that beats HRES across all three energy-critical variables at every lead time, while other AI models fall behind on one or more dimensions. The table below summarises relative performance across the models available on the Jua platform, evaluated against more than 10,000 ground stations using open-source StationBench methodology. All comparisons are like-for-like: station-level RMSE and CRPS, no post-processing, no station fine-tuning. Source: arXiv:2507.09703.
| Model | 10 m Wind / 100 m Wind | 2 m Temperature | Surface Solar Radiation (SSRD) |
|---|---|---|---|
| Jua EPT-2 | Outperforms HRES at every lead time (0–240 h) | Outperforms HRES at every lead time (0–240 h) | Outperforms HRES at every lead time (0–240 h) |
| ECMWF HRES | Benchmark | Benchmark | Benchmark |
| NOAA GFS | Below HRES at all lead times | Below HRES at all lead times | Below HRES at all lead times |
| Microsoft Aurora | Below EPT-2 across full 0–240 h range | Below EPT-2 up to ~130 h lead time | No SSRD output |
| DeepMind GraphCast | Below EPT-2, below HRES on wind | Below EPT-2 | No published SSRD benchmark |
| ECMWF AIFS | Below EPT-2 | Below EPT-2 | Below EPT-2 |
Relative rankings based on StationBench verification (arXiv:2507.09703). “Outperforms” and “below” denote RMSE/CRPS direction versus the stated reference. Aurora has no SSRD output, so EPT-2 wins that variable by default. All EPT-2 results are deterministic flagship; EPT-2e ensemble results are reported separately.
Inside Jua EPT-2: Foundation Model Design And Specs
Jua is a foundation-model and agent company. EPT-2 is the atmospheric fine-tune of the Earth Physics Transformer (EPT), a general spatiotemporal transformer foundation model that learns the governing physics of complex systems, including mass, momentum, and energy conservation, directly from observational data. Jua for Energy is the first applied product built on EPT and Athena, Jua’s AI agent.
EPT-2 operational specifications relevant to energy trading include spatial resolution natively down to 5 km over Europe (EPT-2 HRRR), update frequency up to 24 times per day (EPT-2 RR), and a forecast horizon out to 20 days deterministic and 60 days ensemble. EPT-2e, the ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time, with results documented in arXiv:2507.09703. EPT-2 outperforms leading AI weather models and traditional numerical baselines across all forecast horizons on RMSE, verified at station level with no post-processing.
Aurora and GraphCast run on the Jua platform as guest models, available in the same workspace and benchmarked on the same surface as EPT-2. ECMWF AIFS is also available natively. The comparison comes built into the product.
Frequently Asked Questions
Is ECMWF HRES the most accurate weather model in the world?
ECMWF HRES has been the global benchmark for deterministic medium-range forecasting for four decades and remains the most accurate NWP (numerical weather prediction) model in operational use. As of 2026, Jua’s EPT-2 surpasses HRES across the 0 to 240 hour range on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation, verified against more than 10,000 ground stations using StationBench methodology with no post-processing. EPT-2e, the ensemble variant, also beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time.
Why do energy traders rely on ECMWF rather than GFS?
ECMWF HRES outperforms GFS over European domains at every lead time on the variables that drive energy P&L, including temperature, wind, and precipitation. The structural reasons are resolution at 9 km versus 13 km and 4D-Var data assimilation, which extracts more information from Europe’s dense observational network. ECMWF’s two-week outlook is the reference point for repricing risk around heating demand, renewable output, and system tightness. GFS is most useful as a divergence signal, where disagreement with HRES at day 5–7 quantifies forecast uncertainty before it appears in the price.
What is StationBench and why does it matter for model verification?
StationBench is Jua’s open-source benchmarking methodology that evaluates forecast models against more than 10,000 real ground-station observations globally, with no post-processing or station-specific fine-tuning applied to the model outputs. This approach matters because vendor-provided graphics and gridded verification against reanalysis can mask real-world forecast errors at the locations where energy assets actually operate. Station-level RMSE and CRPS are the metrics that correspond to what a wind farm, solar plant, or gas demand forecast actually experiences.
How does EPT-2e compare to ECMWF ENS for probabilistic forecasting?
EPT-2e, Jua’s ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time, as documented in arXiv:2507.09703. For energy trading, probabilistic skill at day 5–7 is a primary input to risk-management decisions around renewable generation and gas demand. EPT-2e updates 4 times per day and extends to a 60-day horizon. No other AI weather model has shipped a productised ensemble equivalent.
What are the known limitations of ECMWF HRES over Europe?
HRES has documented limitations in three areas relevant to energy markets. First, fog and stratus over the North Sea and English Channel, where boundary-layer parameterisation introduces timing errors in low-cloud onset and dissipation that propagate directly into surface solar radiation forecasts. Second, convective precipitation in summer over the Alps and Balkans, where the 9 km grid cannot explicitly resolve convective initiation. Third, coastal and complex-terrain wind at hub heights, where surface roughness transitions are smoothed at 9 km resolution. EPT-2 HRRR, operating natively down to 5 km over Europe, addresses the resolution component of these limitations.
Conclusion: HRES Remains The Benchmark, EPT-2 Sets The New Bar
ECMWF HRES is the correct benchmark for European medium-range forecasting. Four decades of NWP leadership, 9 km resolution, and 4D-Var data assimilation make it the reference every energy desk runs on. GFS trails HRES over European domains at every lead time on the variables that matter, including temperature, wind, and precipitation. The gap is structural and well-documented.
The 2026 picture introduces a new layer. Jua’s EPT-2, the atmospheric fine-tune of the Earth Physics Transformer general physics foundation model, now holds a verified performance edge over HRES across the tested lead times and variables, with results published in arXiv:2507.09703. EPT-2e beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time. EPT-2 HRRR delivers forecasts natively down to 5 km over Europe, and EPT-2 RR refreshes up to 24 times per day, compared with the 2–4 daily runs that have constrained the energy industry for forty years.
A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves roughly €1.5 M per year. For a multi-GW portfolio, the economics scale linearly. Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec already run Jua for Energy alongside their existing ECMWF subscriptions, using it as the layer that upgrades the forecasting stack around their current feeds.
Jua is a foundation-model and agent company, and Jua for Energy is the first applied product. The verification numbers summarise the case.
Talk to Jua about your portfolio and see the live benchmarks.