Weather Forecasting

2026 Europe AI Weather Benchmarks: EPT-2 vs ECMWF & Aurora

Olivier Lam·June 18, 2026
2026 Europe AI Weather Benchmarks: EPT-2 vs ECMWF & Aurora

Written by: Olivier Lam, Physical AI Team, Jua.ai AG

Key Takeaways for European Energy Traders

  • EPT-2 sets the 2026 accuracy ceiling for AI weather models over Europe. It outperforms ECMWF HRES, AIFS, and Aurora on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation across the full 0–240 hour lead-time range.
  • On StationBench, evaluated against more than 10,000 real ground stations with no post-processing, EPT-2 delivers the lowest RMSE on all four energy-critical variables and beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time.
  • EPT-2’s higher native resolution (~5 km, up to 1 km in Jua for Energy) and training on over 5 petabytes of observational data support stronger performance on European extremes and complex microclimates than earlier AI models.
  • Operationally, EPT-2e runs four times daily and EPT-2 RR up to 24 times daily at a fraction of traditional NWP cost. These accuracy gains translate into millions of euros saved annually on 1 GW wind and solar portfolios.
  • Run these benchmarks on your own region and variables with Jua for Energy.

Europe-Wide RMSE Comparison Across Forecast Horizons

The table below reports RMSE for four energy-critical variables across the full 0–240 hour lead-time range. The evaluation uses the open-source StationBench methodology documented in arXiv:2507.09703 against a global surface-station network, with no post-processing or station fine-tuning. All figures come from the EPT-2 technical report (arXiv:2507.09703).

Model10 m Wind RMSE100 m Wind RMSE2 m Temp RMSESSRD RMSE
EPT-2 (Jua)Lowest across 0–240 hLowest across 0–240 hLowest across 0–240 hLowest across 0–240 h
ECMWF HRESHigher than EPT-2 at all lead timesHigher than EPT-2 at all lead timesHigher than EPT-2 at all lead timesHigher than EPT-2 at all lead times
ECMWF AIFSHigher than EPT-2, operates at ~0.25° gridAdded in AIFS 1.1.0, higher than EPT-2Higher than EPT-2 at all lead timesAdded in AIFS 1.1.0, higher than EPT-2
Microsoft AuroraHigher than EPT-2 across full rangeHigher than EPT-2 across full rangeHigher than EPT-2 up to ~130 hNo SSRD output published

EPT-2 beats ECMWF HRES on every lead time and on all four variables: 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation (SSRD), across the full 0–240 hour range. EPT-2e, the ensemble variant, also beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. Aurora publishes no SSRD output, so EPT-2 is the only AI model in this comparison with full coverage of the four variables that drive a European energy P&L.

Run these benchmarks on your own region and variables on the Jua platform.

Extreme Events and Microclimates in European Energy Markets

The overall RMSE figures above tell only part of the story, because rare extremes drive the largest P&L swings. Extreme-event skill is where the gap between AI model generations matters most for European energy markets. A 2024 Science Advances study evaluating record-breaking extremes in 2018 and 2020 found that earlier-generation AI models GraphCast, Pangu-Weather, and Fuxi systematically underpredicted the intensity and frequency of record-breaking heat, cold, and wind events. Forecast bias grew nearly linearly as record exceedance increased. Those models were trained on 0.25° ERA5 reanalysis data, and ECMWF HRES at 0.1° resolution outperformed all three on 162,751 heat records, 32,991 cold records, and 53,345 wind records from ERA5 in 2020.

EPT-2 represents a structural advance over that generation. It is trained on more than 5 petabytes of observational data from 120+ sources, including proprietary station networks that cover over 10,000 ground stations. This training regime frees EPT-2 from the 0.25° ERA5 distribution that limited its predecessors on out-of-distribution extremes. The architecture learns conservation laws directly from observational data. Outputs are therefore physically constrained by construction rather than extrapolated from a coarse reanalysis grid.

On spatial resolution, EPT-2 HRRR delivers native forecasts at approximately 5 km over Europe. For product deployments inside Jua for Energy, resolution reaches up to 1 km. ECMWF HRES runs at 9 km, and ECMWF AIFS at its N320 operational grid runs at approximately 31 km. That resolution difference is material for complex European terrain. Alpine passes, North Sea coastal gradients, and Iberian plateau microclimates all require sub-10 km resolution to capture the wind and temperature structures that drive generation and demand.

Request a head-to-head comparison on extreme-event skill for your portfolio.

Operational Cadence and Cost for Trading Desks

EPT-2e runs four times per day in operational production, which aligns with existing trading workflows. EPT-2 RR, the rapid-refresh variant, updates up to 24 times per day and keeps intraday traders on fresh numbers. Traditional NWP models such as ECMWF and GFS typically provide four daily runs, with data often 6–12 hours old by the time traders receive it. A typical Jua run completes about 2.5 hours ahead of competing operational runs at the same cycle.

The inference cost differential between AI and traditional NWP spans four orders of magnitude and comes from fundamentally different computational architectures. A single EPT-2 simulation runs on a single GPU in minutes at approximately 0.25 kWh and $0.20–$15. In contrast, a single traditional NWP simulation consumes roughly 8,400 kWh and costs €1,000–€20,000 on HPC infrastructure. Among AI models, AIFS has a comparable inference profile to EPT-2, while EPT-2 runs approximately 25% faster than Aurora at equivalent lead times.

The economic case for accuracy improvement is direct and quantifiable. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves about €1.5 M per year under typical European hedging and imbalance penalty structures. A 1 GW solar portfolio with the same accuracy gain saves roughly €3 M per year. For multi-GW portfolios, these figures scale linearly.

Calculate your potential savings with a custom accuracy analysis.

Accuracy Ranking of European AI Weather Models in 2026

In 2026, EPT-2 is the most accurate AI model over Europe on the StationBench evaluation surface. The benchmark uses the same open-source methodology and station network described above, with no post-processing. EPT-2 outperforms ECMWF HRES on every lead time from 0 to 240 hours on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation. EPT-2e, the ensemble variant, also beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time.

ECMWF AIFS, which became operational in its 1.1.0 version on 27 August 2025, improves on IFS by about one day of forecast skill for precipitation and 12–24 hours for upper-air variables. It still trails EPT-2 on the energy-critical surface variables. Microsoft Aurora trails EPT-2 on 10 m wind, 100 m wind, and 2 m temperature across all forecast horizons and publishes no SSRD output.

Most Accurate Operational Forecast for Europe in 2026

EPT-2, delivered through Jua for Energy, is the most accurate atmospheric forecast in production over Europe in 2026. The evaluation uses open-source StationBench against more than 10,000 real ground stations (arXiv:2507.09703), with no post-processing or station fine-tuning, and any evaluator can reproduce it.

ECMWF HRES remains the universal benchmark and the reference model that serious customers keep in their stack. Jua for Energy does not replace it. ECMWF’s two-week outlook remains the definitive reference point for European traders repricing risk around heating demand, renewable output, and system tightness. EPT-2 now outperforms HRES on the four variables that determine that risk.

The hybrid workflow has become standard. ECMWF provides the raw signal. Jua for Energy handles the plumbing around it and adds EPT-2 as a higher-accuracy layer. This is the production configuration used by Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec.

AI Model Performance on European Record Extremes

Earlier-generation AI models have a documented structural limitation on record extremes. The 2024 Science Advances study showed that GraphCast, Pangu-Weather, and Fuxi, all trained on 0.25° ERA5 data from 1979–2017, systematically underpredicted the intensity and frequency of events outside that training distribution. Bias increased roughly linearly as record exceedance grew. ECMWF HRES at 0.1° resolution outperformed all three on heat, cold, and wind records in both 2018 and 2020 test years.

That study recommended running physics-based NWP and AI models in parallel for high-stakes applications, which remains the correct hybrid-workflow prescription in 2026. EPT-2 advances beyond the 2024 generation on two dimensions that matter for extremes. It is trained on observational data from more than 120 sources, including proprietary station networks, rather than solely on ERA5 reanalysis. Its outputs are physically constrained by conservation laws learned directly from observations, not extrapolated from a coarse grid.

The StationBench evaluation on the same 10,000+ station network, including complex Alpine and coastal terrain, provides the out-of-sample verification that the 2024 study identified as missing from earlier AI model extreme-event claims.

Conclusion: EPT-2 as the 2026 Production Benchmark

The 2026 Europe AI weather benchmark resolves to a clear hierarchy. EPT-2 sets the accuracy ceiling on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation across the full 0–240 hour lead-time range, evaluated on StationBench with no post-processing. EPT-2e beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time. ECMWF AIFS and Microsoft Aurora trail EPT-2 on the energy-critical surface variables, and Aurora publishes no SSRD output.

Operationally, EPT-2e runs four times per day, EPT-2 RR runs up to 24 times per day, and product resolution reaches up to 1 km over Europe inside Jua for Energy. Inference costs approximately $0.20–$15 per simulation on a single GPU, which sits four orders of magnitude below traditional NWP. A four-percentage-point accuracy gain on a 1 GW wind portfolio saves about €1.5 M per year.

Jua is a foundation model and agent company. EPT is a general physics foundation model, and Athena is an AI agent. Jua for Energy is the first applied product. It delivers these benchmarks in production and is already used by European utilities and trading houses executing daily decisions on the platform. The architecture learns physics, and the domain remains a variable.

Schedule a technical evaluation with our forecasting team.

Frequently Asked Questions

How does EPT-2 differ from ECMWF HRES for European energy trading?

ECMWF HRES has been the gold standard in numerical weather prediction for over forty years, running at 9 km horizontal resolution with two to four operational runs per day. EPT-2 outperforms HRES at every forecast horizon on the four variables that drive a European energy P&L, using the same StationBench methodology and station network described above, with no post-processing. EPT-2 HRRR delivers native forecasts at approximately 5 km over Europe, with product resolution up to 1 km inside Jua for Energy.

EPT-2 RR updates up to 24 times per day versus HRES’s two to four. The correct operational posture is augmentation rather than replacement. Serious customers keep their ECMWF subscription and run Jua for Energy alongside it. Jua for Energy displaces the plumbing around the ECMWF feed, including the grib pipeline, manual benchmarking, and morning-briefing assembly, while preserving the underlying signal.

What is StationBench and why does it matter for European forecast evaluation?

StationBench is Jua’s open-source benchmarking methodology that evaluates forecast skill against real ground-truth observations from more than 10,000 surface stations globally, with no post-processing or station fine-tuning applied to any model. This matters for European energy applications because it measures what traders and meteorologists actually care about: how accurately a model predicts conditions at specific locations such as wind farm sites, solar parks, and demand centers, rather than on a coarse reanalysis grid.

Earlier AI model evaluations that relied solely on ERA5 reanalysis as ground truth systematically overstated skill on extremes, because ERA5 smooths out the intensity of record events. StationBench removes that artifact. The EPT-2 technical report (arXiv:2507.09703) documents the full methodology and results, and the benchmarking surface inside Jua for Energy allows any customer to run the same comparison on their own region and variable in under five minutes.

Can AI models be trusted for European extreme weather events relevant to energy markets?

Trust depends on the model generation and training regime. Earlier AI models GraphCast, Pangu-Weather, and Fuxi were documented in a 2024 Science Advances study to systematically underpredict the intensity and frequency of record-breaking heat, cold, and wind events, with bias growing linearly as record exceedance increased. That limitation was structural because those models were trained on 0.25° ERA5 reanalysis data from 1979 to 2017 and could not reliably extrapolate to events outside that distribution.

EPT-2 addresses this at the architecture level. It is trained on more than 5 petabytes of observational data from over 120 sources, including proprietary station networks, rather than solely on ERA5. Its outputs are physically constrained by conservation laws learned directly from observations. The StationBench evaluation on the same 10,000+ station network, including Alpine, coastal, and other complex European terrain, provides the out-of-sample verification that earlier AI models lacked.

The recommended workflow for high-stakes extreme-event applications remains a hybrid one. EPT-2 serves as the primary AI signal, ECMWF HRES remains the physics-based reference, and Jua for Energy’s divergence alerts fire the moment the two disagree on a key variable.

How does Jua for Energy’s refresh frequency affect intraday trading decisions?

Traditional NWP models deliver four global forecasts per 24-hour period, with data typically six to twelve hours old by the time it reaches a trader’s screen. Between runs, traders operate on stale numbers. EPT-2 RR updates up to 24 times per day, with a typical run completing about 2.5 hours ahead of competing operational runs at the same cycle.

Actual-generation power forecasts inside Jua for Energy refresh every 15 minutes. For intraday trading in European short-term power markets, where the 0–48 hour window has the largest P&L impact from forecast freshness, this cadence difference is the primary operational advantage. Jua for Energy’s correction alerts fire the moment a model revises its own output between runs, and divergence alerts fire the moment two models disagree on a key variable. Both alert types surface trade windows before the market reprices.

How quickly can a quant team or meteorologist evaluate EPT-2 against their current provider?

The live benchmark on the Jua platform returns a head-to-head accuracy comparison in under five minutes. A meteorologist or quant developer selects a region, a variable, and a time window, adds their current provider alongside EPT-2, and the platform returns RMSE and CRPS results against StationBench ground truth on the spot.

Backtests against years of historical forecasts run in about five minutes via Athena, Jua’s AI agent. For teams that prefer programmatic access, the Python SDK installs with pip install jua, and the REST API exposes 25-plus models, including ECMWF HRES, ECMWF AIFS, Aurora, and GraphCast, through a single schema with Apache Arrow support for large payloads. Hindcast data is available across multiple Jua and third-party models for strategy backtesting. The evaluation that takes a quant team a quarter to build elsewhere stands up in days.

View the key takeaways as a web story

Want to talk to the team behind the writing?

Book a demo to see EPT-2 and Athena in production, or read the open papers behind the work.