Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways for European Energy Desks
- EPT-2 now leads ECMWF HRES on every lead time (0–240 h) for 10 m wind, 100 m wind, 2 m temperature, and SSRD across Europe.
- The model is a physics-constrained spatiotemporal transformer trained on 5+ PB of observational data, which removes the extrapolation failures seen in other AI weather models.
- At ~5 km resolution, EPT-2 runs in minutes on a single GPU for ~$0.20–$15, compared with thousands of euros and hours for traditional NWP on HPC.
- Energy-trading portfolios gain about €1.5 M per GW of wind and €3 M per GW of solar for every four-percentage-point accuracy improvement.
- Start your free benchmark with Jua to compare EPT-2 against your current provider on your own region and variables in under five minutes.
EPT-2 vs ECMWF HRES: New Accuracy Leader for Europe
ECMWF HRES has led numerical weather prediction (NWP) for over forty years and remains the definitive reference point for European energy traders repricing risk around heating demand, renewable output, and system tightness. That benchmark has now been surpassed.
EPT-2 outperforms ECMWF HRES on every lead time across the full 0–240 hour range on 10 m wind, 100 m wind, 2 m temperature, and SSRD. The evaluation methodology is open-source StationBench, verified against more than 10,000 real ground stations with no post-processing or station fine-tuning. This standard makes the result auditable by any meteorologist running their own benchmark.
EPT-2 is a general spatiotemporal transformer foundation model, the Earth Physics Transformer, trained on 5+ petabytes of observational data from more than 120 distinct sources. Its outputs are physically constrained by construction. The architecture learns conservation laws governing mass, momentum, and energy directly from observational data, in a latent representation integrated forward in time. This structure explains why EPT-2 does not exhibit the extrapolation failures that a Science Advances analysis identified in GraphCast, Pangu-Weather, and FuXi on record-breaking temperature and wind events. Those models lack explicit physical constraints, while EPT-2 enforces them.
Beyond accuracy, EPT-2 delivers operational advantages in both resolution and cost. EPT-2 runs at up to ~5 km spatial resolution over Europe (EPT-2 HRRR variant), compared with 9 km for ECMWF HRES. A single EPT-2 inference runs on a single GPU in minutes at approximately 0.25 kWh and $0.20–$15. An equivalent NWP simulation on HPC consumes approximately 8,400 kWh and €1,000–€20,000. The cost asymmetry is roughly four orders of magnitude.
Run a live comparison to see EPT-2 head-to-head against your current forecast provider on your own region and variable in under five minutes.
Regional Models in Europe: ECMWF vs ICON and AROME
ICON-EU, operated by Germany’s DWD, runs at higher resolution than ICON Global over its European domain and retains a performance window at short lead times (0–48 hours) for convective-scale events and orographic wind in Central Europe. KNMI HARMONIE-AROME and Météo-France AROME provide comparable short-range skill in their respective national domains, with mesh spacing that resolves coastal and complex-terrain effects that global models smooth over.
These regional models are operationally valuable for intraday and day-ahead decisions in their native domains. Their limitations are well-defined. Update frequency is typically two to four runs per day, forecast horizons extend to roughly 60–72 hours, and cross-border coverage degrades at domain edges. For medium-range (3–10 day) positioning, the window where most European power and gas trading risk is priced, global deterministic models dominate, and EPT-2 leads that tier. All of these models, including ICON-EU, ICON Global, and HARMONIE-AROME, are available on the Jua platform alongside EPT-2 under a unified schema, so comparison requires no pipeline re-engineering.
EPT-2e Ensembles: Managing 3–15 Day Uncertainty
Deterministic forecasts degrade beyond roughly five days as atmospheric predictability limits are reached. Ensemble forecasting, which runs multiple perturbed simulations to sample forecast uncertainty, is the standard method for quantifying that uncertainty and pricing probabilistic risk.
EPT-2e, the ensemble variant of EPT-2, beats the 50-member ECMWF ENS mean on both RMSE and CRPS (Continuous Ranked Probability Score) at virtually every lead time. CRPS is the standard metric for ensemble calibration. It penalises both sharpness and reliability simultaneously, which makes it the correct measure for probabilistic forecast value in trading applications. EPT-2e updates four times per day with a ~10-day horizon, compared with 15 days for ECMWF ENS.
For European energy traders, ensemble skill at 3–15 days translates directly to better-calibrated positions on week-ahead gas demand, wind-generation spread, and temperature-driven load. A sharper, better-calibrated ensemble means tighter confidence intervals on the variables that move the spread and earlier visibility on the tail scenarios that move it most.
Deterministic Accuracy Rankings for Europe (0–240 h)
| Model | 10 m Wind | 100 m Wind | 2 m Temperature | SSRD |
|---|---|---|---|---|
| EPT-2 | #1 across full 0–240 h range | #1 across full 0–240 h range | #1 across full 0–240 h range | #1 across full 0–240 h range |
| ECMWF HRES | #2 — 40-year NWP benchmark | #2 — 40-year NWP benchmark | #2 — 40-year NWP benchmark | #2 — 40-year NWP benchmark |
| Microsoft Aurora | #3 — loses to EPT-2 across full range | #3 — loses to EPT-2 across full range | #3 — loses to EPT-2 up to ~130 h | No SSRD output |
| ICON-EU | Short-range regional advantage (0–48 h, native domain) | Short-range regional advantage (0–48 h, native domain) | Short-range regional advantage (0–48 h, native domain) | Limited medium-range skill vs EPT-2 |
Aurora’s fixed 6-hour roll-forward architecture compounds error at longer lead times, while EPT-2 forecasts at native any-Δt without rolling. Aurora also required 32 × A100 GPUs over 18 days to train. EPT-2 was trained on 8 × H100 GPUs in 10 days, using four times fewer GPUs and a substantially shorter training cycle.
Energy-Trading P&L Impact of Higher Forecast Accuracy
Forecast accuracy directly affects P&L in European power markets. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 M per year, with the figure scaling linearly across multi-GW portfolios.
The market-sizing economics are concrete. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 M per year under typical hedging and imbalance-penalty structures. A 1 GW solar portfolio at the same accuracy gain saves approximately €3 M per year. The solar figure is larger because SSRD forecast error maps directly to day-ahead generation shortfalls, which clear at intraday premium prices in tight European markets.
European energy traders are already turning to AI tools to gain an edge on forecast revisions, so the decision now focuses on which model delivers the accuracy that moves the P&L. EPT-2’s lead over ECMWF HRES on SSRD matters especially for solar portfolios. Aurora produces no SSRD output at all, which leaves EPT-2 as the only AI model in the comparison set with a complete variable surface for European renewables trading.
Customers operating multi-GW portfolios, including some of Europe’s largest energy companies, commodity traders, and hedge funds such as Axpo, TotalEnergies, Statkraft, EnBW, and EDF, scale these economics linearly. At 5 GW of wind exposure, a four-percentage-point accuracy gain is worth approximately €7.5 M per year.
Live Benchmarking and Alerts on the Jua Platform
Jua for Energy exposes more than 25 models on a single platform, including 10 proprietary AI models from the EPT family plus 15 third-party NWP and AI models such as ECMWF HRES, ECMWF ENS, ECMWF AIFS, NOAA GFS, GFS GraphCast, Microsoft Aurora, DWD ICON Global, and ICON-EU. Any meteorologist or quant developer can select a region, a variable, and a time window and run a head-to-head benchmark in under 30 seconds.
The benchmarking surface is the deal trigger for most Jua for Energy evaluations. Meteorologists who arrive sceptical of vendor accuracy claims become internal champions the moment they run the comparison on their own highest-stakes region, because the numbers speak for themselves. This conversion happens quickly because Athena, Jua’s AI agent instrumented with the Jua for Energy tool surface, turns a natural-language question into a briefing, a benchmark, a backtest, or a custom widget in approximately 90 seconds, or a full backtest in approximately five minutes. Sales cycles at Jua have compressed to as little as two weeks for technically-led evaluations, because the proof runs live on the prospect’s own data.
EPT-2 RR updates up to 24 times per day, and EPT-2 HRRR delivers the same cadence at up to ~5 km spatial resolution over Europe. Traditional NWP providers refresh two to four times per day. Between runs, traders on legacy stacks are looking at stale numbers. Divergence alerts on the Jua platform fire the moment two models disagree on a key variable, and correction alerts fire the moment a model revises its own output. You act before the market does.
See a guided walkthrough to watch EPT-2 benchmarked against ECMWF HRES and Aurora on your region and variable, live, in under five minutes.
Frequently Asked Questions
Is EPT-2 a replacement for ECMWF HRES in a European energy-trading workflow?
No, and Jua for Energy is not designed as a replacement. Most serious customers keep their ECMWF subscription and run Jua for Energy alongside it, and ECMWF AIFS, ECMWF’s own AI model, even runs natively on the Jua platform. EPT-2 and the Jua for Energy product surface replace the plumbing around the ECMWF feed: the in-house grib pipeline, the manual benchmarking harness, the morning-briefing analyst, and the spreadsheet stitching.
The 7–9 a.m. manual prep routine compresses into a single workspace, refreshed up to 24 times per day, where every model, including ECMWF HRES, ENS, AIFS, Aurora, and EPT-2, runs under one schema and one API. EPT-2 outperforms ECMWF HRES on every lead time and variable, but the value proposition is the complete stack, not a one-for-one swap.
How does EPT-2 handle surface solar radiation forecasting compared to other AI weather models?
EPT-2 produces SSRD output natively and leads ECMWF HRES on SSRD accuracy across the full 0–240 hour range, verified against more than 10,000 ground stations on open-source StationBench. Microsoft Aurora produces no SSRD output at all, which means it cannot be used directly for solar generation forecasting without a separate radiation model.
For European solar portfolios, where day-ahead generation shortfalls clear at intraday premium prices, this gap is commercially significant. A 1 GW solar portfolio that gains four percentage points of SSRD forecast accuracy saves approximately €3 M per year under typical European market structures. EPT-2 is currently the only AI model in the leading comparison set with a complete variable surface for European renewables trading, including SSRD.
What is the difference between EPT-2 and EPT-2e, and when should each be used?
EPT-2 is the deterministic flagship, a single best-estimate forecast, updated four times per day, running out to 20 days at up to ~5 km spatial resolution over Europe. It is the right tool for day-ahead and intraday positioning where a single high-accuracy forecast drives the decision.
EPT-2e is the ensemble variant. It outperforms ECMWF ENS on the metrics that matter for probabilistic forecasting, RMSE and CRPS, as detailed earlier, updates four times per day, and extends to a ~10-day horizon. It is the right tool for medium-range (3–15 day) uncertainty quantification, including week-ahead gas demand spreads, wind-generation confidence intervals, and temperature tail scenarios for heating demand. In practice, most European energy desks use both: EPT-2 for near-term deterministic positioning and EPT-2e for probabilistic risk management at longer horizons. Both are accessible through the Jua platform’s API, Python SDK, and Athena agent surface under a unified schema.
How quickly can a quant team integrate EPT-2 into an existing trading pipeline?
The Python SDK installs with a single command: pip install jua. The REST API exposes more than 25 models, including EPT-2, EPT-2e, ECMWF HRES, Aurora, and GraphCast, through a unified schema with Apache Arrow support for large payloads.
Hindcast data is available across multiple Jua and third-party models for backtesting. A backtest that would take a quant team a quarter to build from raw model subscriptions runs in approximately five minutes via Athena, or directly through the SDK for teams that prefer programmatic access. ENTSO-E grid data integrates natively for European power-market context. Documentation is at docs.jua.ai, and the developer dashboard is at developer.jua.ai.
Conclusion: EPT-2 at the Top of the European Forecast Hierarchy
The deterministic hierarchy for European weather forecasting now has a new leader. As of June 2026, EPT-2, the flagship model inside Jua for Energy and the first applied product from Jua’s foundation model and agent platform, leads the deterministic hierarchy across all four variables and the full forecast horizon documented in the benchmark. EPT-2e leads the 50-member ECMWF ENS on RMSE and CRPS at virtually every lead time. Microsoft Aurora ranks third on wind and temperature and produces no SSRD output. ICON-EU retains short-range regional value in its native domain.
The P&L translation is direct, at €1.5 M per GW of wind and €3 M per GW of solar at four percentage points of accuracy gain. For multi-GW portfolios, the economics scale linearly. The proof runs live on the Jua platform in under five minutes, on any region and variable a trading desk cares about.
Jua is a foundation model and agent company. EPT is a general physics foundation model, and Athena is an AI agent. The atmosphere is the first physical system EPT has been fine-tuned for. Energy trading is the first market Athena has been instrumented for. Both will expand.
Schedule a walkthrough to see the June 2026 benchmark hierarchy live on your region, your variable, and your current forecast provider.
