Weather Forecasting

AI Hourly Weather Model RMSE: EPT-2 Beats ECMWF & GraphCast

Olivier Lam·May 18, 2026
AI Hourly Weather Model RMSE: EPT-2 Beats ECMWF & GraphCast

Written by: Olivier Lam, Physical AI Team, Jua.ai AG | Last updated: June 26, 2026

Key Takeaways for Energy Traders

  • RMSE is the decisive metric for energy traders at short lead times, where small temperature or wind errors can move spreads and positions before intraday gate closure.
  • EPT-2 and EPT-2e set the current StationBench standard at 1 h, 3 h, and 6 h lead times for 2 m temperature, 100 m wind speed, and precipitation across more than 10,000 ground-truth stations.
  • The native any-Δt architecture lets EPT-2 output forecasts at arbitrary time steps without interpolation or rolling, so its largest advantage appears at the 1 h and 3 h windows that matter most for intraday power markets.
  • Improved hourly accuracy translates directly into P&L: a 1 GW wind portfolio gaining four percentage points of forecast skill can save approximately €1.5 M per year, with similar gains for solar portfolios.
  • See EPT-2 on your assets to run live StationBench comparisons on your own region and variables and see the P&L delta in under five minutes.

How StationBench Compares AI and NWP Models

StationBench reports RMSE across five models and three variables at each lead time. Every figure is station-verified against more than 10,000 ground observations, and lower RMSE is better. ECMWF HRES is the NWP benchmark, while GraphCast (Google DeepMind) and Aurora (Microsoft) are leading AI peers.

See EPT-2 vs. your current provider to compare models on your own region and variables.

1-Hour Lead-Time RMSE: Shortest Horizon, Largest Edge

At 1 h lead time, EPT-2 and EPT-2e both outperform ECMWF HRES, GraphCast, and Aurora on all three variables evaluated by StationBench across more than 10,000 stations. EPT-2’s native any-Δt architecture, trained to predict at arbitrary time steps rather than rolling forward in fixed 6-hour increments, is the structural reason for the 1 h advantage. Aurora and GraphCast must interpolate from a 6-hour grid, which compounds error at sub-6-hour lead times.

The 1 h StationBench results show that EPT-2 achieves the lowest RMSE across 2 m temperature, 100 m wind speed, and precipitation, with the largest relative gain in wind speed where interpolation artifacts are most pronounced.

Source: StationBench. Station count: >10,000. No post-processing applied.

3-Hour Lead-Time RMSE: Intraday Trading Window

At 3 h, the EPT family widens its lead over NWP and AI peers. The any-Δt advantage described above extends directly to this 3 h window. ECMWF HRES, operating on a fixed output grid, cannot natively resolve 3 h outputs without interpolation, while EPT-2 produces 3 h forecasts as direct model outputs.

StationBench confirms that EPT-2 and EPT-2e maintain strong performance across all three variables at this lead time. This 3 h horizon maps directly to intraday power market positions in Germany, Great Britain, France, the Netherlands, and Belgium, the five countries where Jua for Energy delivers live power forecasts.

Source: StationBench. Station count: >10,000. No post-processing applied.

EPT-2e, the ensemble variant, performs competitively with the ECMWF ENS mean on both RMSE and CRPS at short lead times.

6-Hour Lead-Time RMSE: Native Grid for Fixed-Step Models

At 6 h, the fixed-step architecture of Aurora and GraphCast produces its first native output. Their 1 h and 3 h results are interpolated, while the 6 h result is the first point where those models operate at their trained resolution. EPT-2 performs well at 6 h on all three variables and keeps the any-Δt advantage.

The native any-Δt design continues to matter at this horizon. EPT-2 does not roll forward in fixed steps and does not accumulate the interpolation error that Aurora and GraphCast carry into every sub-6-hour output.

Source: StationBench. Station count: >10,000. No post-processing applied.

EPT-2 produces forecasts natively at high spatial resolution over Europe. That spatial granularity can matter at 6 h, because a wind ramp localized to a coastal corridor can be resolved better by higher resolution models and missed by coarser grids.

Extreme-Event RMSE and Alerting in Jua for Energy

Standard RMSE averages across all conditions. StationBench’s extreme-event evaluation isolates the top and bottom deciles of observed values, which captures the wind ramps, cold snaps, and convective precipitation events that drive the largest P&L swings in energy markets. EPT-2 and EPT-2e maintain an advantage over ECMWF HRES, GraphCast, and Aurora in this extreme-event subset.

That advantage becomes actionable through Jua for Energy’s alert system, which is designed to surface extreme events as they develop. Divergence alerts fire the moment two or more models disagree on a key variable, which signals that an extreme event is in play and that the models have not converged. Correction alerts fire the moment a model revises its own output between runs. Both alert types are filterable by zone and by PSR (Production Source Resource) type. A wind ramp that EPT-2 resolves at 1 h and that ECMWF HRES misses at 6 h creates a trade window for the desk that sees it first.

Extreme events evolve rapidly, so the model that updates most frequently has the first chance to resolve them. EPT-2 HRRR updates 4x/day over Europe at high spatial resolution, against the two to four daily runs that traditional NWP infrastructure can sustain. The compute economics explain why: a single EPT-2 inference runs at approximately 0.25 kWh and $0.20–$15 on a single GPU, versus approximately 8,400 kWh and €1,000–€20,000 per NWP simulation on HPC. Roughly four orders of magnitude cheaper at run time makes the refresh cadence that was physically impossible for NWP operationally routine for EPT-2, and that higher refresh rate translates into earlier detection of the extreme events that drive the largest P&L swings.

Energy-Trading P&L Implications of Better RMSE

Forecast accuracy translates directly to P&L through imbalance costs and hedging efficiency. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 M per year under typical European hedging and penalty structures. A 1 GW solar portfolio at the same accuracy gain saves approximately €3 M per year, and customers operating multi-GW portfolios scale these economics linearly.

The RMSE improvements described above translate directly into portfolio-level savings. For a trading house running 5 GW of wind across Germany and Great Britain, switching from ECMWF HRES to EPT-2 at short lead times could save approximately €7.5 M per year, scaling the €1.5 M per GW figure established earlier.

Athena, Jua’s AI agent instrumented with the Jua for Energy tool surface, turns that analysis into a deliverable in natural language. A trader can ask: “What is the EPT-2e versus ECMWF ENS wind spread for northern Germany over the next six hours, and what does the divergence imply for the intraday position?” Athena returns the answer, the underlying widget, and the model delta in approximately 90 seconds. The 7–9 a.m. manual prep routine compresses into a single workspace open before the market does.

Calculate your portfolio’s P&L delta to run a live StationBench comparison on your own region and variables and see the savings in under 5 minutes.

Run These Benchmarks Yourself on the Jua Platform

The StationBench numbers in this article are reproducible on the Jua platform in real time. The live benchmarking surface puts 25+ models on a single workspace, including 10 proprietary AI models from the EPT family plus 15 third-party NWP and AI models such as ECMWF HRES, ECMWF ENS, ECMWF AIFS, NOAA GFS, GraphCast, and Aurora, on any region, any variable, and any time window. A head-to-head result returns in seconds. Meteorologists who were sceptical of vendor accuracy claims often become internal champions once they run the benchmark themselves.

Quant developers can pipe the same models directly into their own systems via pip install jua. The REST API exposes all 25+ models through a single schema with Apache Arrow support for large payloads. Hindcast data is available across multiple Jua and third-party models for backtesting. A backtest that takes a quant team a quarter to build elsewhere runs in approximately 5 minutes via Athena.

Jua is a foundation model and agent company, and Jua for Energy is the first applied product. EPT is a general physics foundation model, so the same architecture that learns atmospheric dynamics already predicts plasma behaviour inside a tokamak. Athena is an AI agent, currently instrumented with the Jua for Energy tool surface. The atmosphere is the first physical system, and energy trading is the first market, and both will expand.

Explore the Jua platform to see EPT-2 and EPT-2e against your current provider.

Frequently Asked Questions

What is StationBench and why does it matter for evaluating AI hourly weather model RMSE?

StationBench is Jua’s open-source evaluation framework that measures forecast RMSE against more than 10,000 real ground-truth weather stations globally, with no post-processing or station fine-tuning applied. Most AI weather model benchmarks use gridded reanalysis data, such as ERA5, as the verification target. Reanalysis is itself a model output, which means benchmarking against it measures how well one model agrees with another, not how well it predicts what actually happened at a physical location.

StationBench eliminates that circularity. Every RMSE figure in this article is verified against observed station data, which makes the numbers directly comparable to the ground-truth conditions that energy traders, meteorologists, and quant developers care about. These conditions include what the wind actually did at a turbine hub height, what the temperature actually was at a load centre, and what precipitation actually fell over a hydro catchment. The methodology is documented in the StationBench repository.

What is the native any-Δt advantage and why does it matter at 1 h and 3 h lead times?

Most AI weather models, including Aurora and GraphCast, are trained on a fixed 6-hour output grid. To produce a 1 h or 3 h forecast, they must either interpolate between 6-hour steps or roll the model forward in 6-hour increments and extract intermediate states. Both approaches compound error, because interpolation introduces smoothing artifacts and rolling forward accumulates the model’s own error at each step.

EPT-2 is trained to predict at arbitrary time steps, any Δt, without rolling. It produces a 1 h forecast as a direct model output, not as a derived approximation. The RMSE results described above reflect this structural difference. EPT-2’s advantage over Aurora and GraphCast is largest at 1 h, where those models are furthest from their native resolution, and narrows slightly at 6 h, where they operate at their trained grid. For intraday energy trading, where the 1 h and 3 h windows map directly to gate-closure and balancing-market positions, the native any-Δt architecture separates a model designed for hourly forecasting from one that was not.

How does EPT-2e’s ensemble performance compare to ECMWF ENS at short lead times?

EPT-2e is the ensemble variant of EPT-2 and significantly surpasses the ECMWF ENS mean for medium- to long-range forecasting. CRPS measures the full probabilistic skill of an ensemble, including both mean error and the calibration of the spread, which makes it the standard metric for evaluating whether an ensemble’s uncertainty estimates are trustworthy. EPT-2e achieves this with 10 published members against ECMWF ENS’s 50, which reflects the efficiency of the EPT architecture rather than a limitation of the ensemble.

For energy traders who position around probabilistic forecasts, such as wind-ramp probability, cold-snap tail risk, and precipitation exceedance thresholds, EPT-2e provides ensemble skill that can exceed the NWP gold standard at certain lead times where intraday and day-ahead decisions are made. Like EPT-2, EPT-2e updates 4x/day on the Jua platform, so ensemble forecasts maintain the same high refresh cadence.

How does improved hourly RMSE translate into energy-trading P&L?

Improved hourly RMSE affects P&L through imbalance costs and hedging efficiency. In European power markets, a balancing-responsible party (BRP) that submits a generation schedule and then deviates from it pays imbalance penalties, and the size of the penalty depends on how far the forecast was from actual generation. A lower RMSE at 1 h, 3 h, and 6 h lead times keeps the submitted schedule closer to what the asset actually produces, which reduces imbalance exposure directly.

The market-sizing economics are concrete. As noted earlier, a 1 GW wind portfolio saves approximately €1.5 M per year and a 1 GW solar portfolio saves approximately €3 M per year at a four-percentage-point accuracy gain. These figures scale linearly with portfolio size. For a trading house or utility running 5–10 GW of renewables, the annual P&L implication of switching from ECMWF HRES to EPT-2 at short lead times sits in the range of €7.5 M to €30 M, depending on the mix of wind and solar and the penalty regime of the markets traded.

Conclusion: Why EPT-2 Matters for Intraday Power

The StationBench evaluation establishes EPT-2 and EPT-2e as strong performers for AI hourly weather models at short lead times across 2 m temperature, 100 m wind speed, and precipitation, verified against more than 10,000 ground-truth stations. Across standard and extreme-event conditions, EPT-2 and EPT-2e show consistent advantages over leading NWP and AI models at the short lead times that drive intraday trading decisions.

For a 1 GW wind portfolio, improved forecast accuracy can translate to significant savings, and similar gains apply for a 1 GW solar portfolio. The live benchmark on the Jua platform returns a head-to-head result on your own region and variables in seconds, so the numbers remain transparent and reproducible.

Run StationBench on your region to compare EPT-2 with your current provider and act before the market does.

View the key takeaways as a web story

Want to talk to the team behind the writing?

Book a demo to see EPT-2 and Athena in production, or read the open papers behind the work.