New 2026 Benchmarks Show EPT-2 Leading AI Weather Models

AI Weather Forecasting: 2026 Guide to Top Models

ON THIS PAGE

Written by: Olivier Lam, Physical AI Team, Jua.ai AG | Last updated: June 24, 2026

Key Takeaways for Energy and Weather Teams

  • Physics-constrained AI models like EPT-2 learn conservation laws from data and avoid the physically impossible outputs that generic machine-learning forecasts produce.
  • June 2026 benchmarks show EPT-2 outperforming ECMWF HRES on RMSE for wind, temperature, and solar radiation across all 0–240 hour lead times.
  • EPT-2e’s 30-member ensemble beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time, evaluated on over 10,000 global stations.
  • Native any-Δt forecasting and single-GPU inference make EPT-2 roughly 25% faster and four orders of magnitude cheaper to run than traditional NWP or competing AI models.
  • Energy traders can benchmark EPT-2 and 25+ models on their own region and variables by booking a demo with Jua.

Latest Benchmark Releases for 2026 Performance

Two technical reports define the current state of AI weather forecasting accuracy. arXiv:2507.09703 documents EPT-2, Jua’s flagship deterministic model. arXiv:2410.15076 documents EPT-1.5, its predecessor. Together they establish a continuous performance record against the global benchmark set.

EPT-2 outperforms ECMWF HRES on every lead time across the four variables that drive energy P&L: 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation (SSRD), across the full 0–240 hour range. EPT-2e, the ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE (root mean square error) and CRPS (continuous ranked probability score) at virtually every lead time. These results are evaluated against more than 10,000 real ground stations using the open-source StationBench methodology, with no post-processing or station fine-tuning applied.

These benchmark results come from Jua’s foundation model research, which the company has turned into products for energy markets. Jua for Energy represents the first commercial application of the EPT and Athena models, in a foundation-to-product relationship similar to how Anthropic built Claude Code on top of its core Claude models.

See EPT-2 benchmarked against your current provider and quantify the gap on your own assets.

StationBench Methodology and AI Forecast Accuracy

StationBench provides the independent framework that validates these AI weather results. It is an open-source evaluation system that measures forecast skill against observations from more than 10,000 ground stations globally, without post-processing or station-specific fine-tuning. This methodology underpins all published EPT-family benchmark claims.

The StationBench results confirm EPT-2’s advantage across all four energy-critical variables mentioned above. EPT-2e, with 30 ensemble members, beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time, and does so with fewer members. EPT-1.5 outperforms GraphCast, FuXi, Pangu-Weather, and ECMWF HRES on European wind and temperature.

Traders and meteorologists can run benchmarks on their own region and variables on the Jua platform at athena.jua.ai.

EPT-2 vs Aurora: Head-to-Head Performance for Energy Variables

EPT-2 beats Microsoft Aurora on 10 m wind, 100 m wind, and 2 m temperature across the full 0–240 hour range. On SSRD, EPT-2 wins by default because Aurora produces no SSRD output.

The architectural difference between the models drives this gap. EPT-2 uses native any-Δt forecasting and is trained to predict at arbitrary lead times rather than rolling forward in fixed 6-hour increments. Aurora and most peer models roll forward in 6-hour steps, which compounds error at each step. EPT-2 does not roll. EPT-2 inference runs approximately 25% faster than Aurora, on a single GPU, at roughly 0.25 kWh and $0.20–$15 per simulation. Aurora required 32 × A100 GPUs over 18 days to train, while EPT-2 trained on 8 × H100 GPUs in 10 days.

On the ensemble side, EPT-2e ships as a productised 30-member ensemble that beats the 50-member ECMWF ENS mean on RMSE and CRPS. Aurora has no productised ensemble equivalent. Aurora and GraphCast remain research outputs from large AI labs, while Jua for Energy operates as a production platform where both models run as guests on the same benchmarking surface.

Five-Model Comparison: Best AI Weather Models in 2026

Model Deterministic RMSE vs HRES (10 m wind, 100 m wind, 2 m temp, SSRD, 0–240 h) Ensemble / CRPS Update Frequency
EPT-2 Beats HRES on every lead time across all four variables EPT-2e (30 members) beats 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time Up to 24×/day (EPT-2 RR); EPT-2e 4×/day
EPT-2e Ensemble mean beats HRES on RMSE at virtually every lead time 30 members; beats 50-member ENS mean on CRPS at virtually every lead time 4×/day
Microsoft Aurora Loses to EPT-2 on 10 m wind, 100 m wind across full range; loses on 2 m temp up to ~130 h; no SSRD output No productised ensemble Typically 4×/day (research cadence; no productised operational schedule)
GraphCast (GFS-init) Loses to EPT-1.5 on European wind and temperature No productised ensemble Typically 4×/day (research cadence)
ECMWF HRES / ENS The 40-year benchmark, and universal reference for deterministic NWP skill ENS: 50 members; gold standard for probabilistic NWP 2–4×/day

Why Physics-Based Foundation Models Beat Generic AI

Physics-constrained transformers avoid the hallucination problems that affect generic AI models on atmospheric data. Standard transformers applied naively to weather fields can violate conservation laws and produce impossible pressure gradients or energy budgets. EPT solves this by learning physical constraints directly from observational data rather than relying on symbolic rules.

The Earth Physics Transformer is a general spatiotemporal transformer foundation model that learns the governing physics of complex systems in a latent representation integrated forward in time. The architecture is constrained by construction. Outputs respect mass, momentum, and energy conservation because the model internalises those constraints from the data.

The compute asymmetry between traditional NWP and EPT-2 is equally significant. A single traditional NWP simulation consumes approximately 8,400 kWh and costs €1,000–€20,000 on HPC infrastructure, running for one to two hours. A single EPT-2 inference runs on a single GPU in minutes at roughly 0.25 kWh and $0.20–$15. The four-orders-of-magnitude cost advantage described in the Aurora comparison above is what makes 24 daily refreshes operationally viable where traditional NWP is capped at two to four.

The same benchmarking tools are available at athena.jua.ai for custom regional analysis.

Operational Advantages for Energy Traders

Jua’s forecasts translate benchmark gains into portfolio-level P&L impact in European energy markets. Improved forecast accuracy for wind and solar assets compounds across fleets and can yield material annual savings that scale with portfolio size.

EPT-2 RR updates up to 24 times per day, compared with the two-to-four daily runs that define traditional NWP. EPT-2 HRRR delivers high-resolution coverage at about 5 km resolution over Europe. Actual-generation power forecasts refresh every 15 minutes with a 48-hour horizon.

Athena, Jua’s AI agent instrumented with the Jua for Energy tool surface, resolves a typical natural-language query in approximately 90 seconds. Morning briefings, model divergence checks, and custom widgets all run through the same interface. Backtests complete in approximately 5 minutes. Jua serves major utilities across five continents, including some of Europe’s largest energy companies, as well as commodity traders and hedge funds. Customers include Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec.

Integration for Quants and Developers: pip install jua

Jua for Energy exposes 25+ models through a single REST API (POST /v1/forecast/data) with Apache Arrow support for large payloads. The Python SDK installs via pip install jua from PyPI and provides forecast access, hindcast and backtesting, and weather-parameter standardisation across all models.

The 25+ models on the platform include EPT-2, EPT-2e, EPT-2 RR, EPT-2 HRRR, Microsoft Aurora, GFS GraphCast, ECMWF AIFS, ECMWF HRES, ECMWF ENS, NOAA GFS, DWD ICON Global, DWD ICON-EU, and others, all under a unified schema. Swapping or comparing models requires no pipeline re-engineering. ENTSO-E grid data integrates directly for European power-market context. Documentation is at docs.jua.ai and the developer dashboard at developer.jua.ai.

Schedule a technical integration walkthrough to review APIs, schemas, and deployment options with Jua’s team.

Frequently Asked Questions

How accurate is AI weather forecasting compared to ECMWF HRES in 2026?

EPT-2, Jua’s flagship deterministic model, outperforms ECMWF HRES on every lead time from 0 to 240 hours across the four variables most relevant to energy trading: 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation. This is measured on StationBench against more than 10,000 real ground stations with no post-processing applied. EPT-2e, the ensemble variant with 30 members, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. ECMWF HRES remains the universal benchmark for traditional NWP and the reference against which all AI models are evaluated, and Jua for Energy runs alongside it rather than replacing it.

What is the best AI weather model for energy trading in 2026?

On published StationBench benchmarks, EPT-2 leads all evaluated models on the variables that drive energy P&L. For probabilistic and ensemble applications, EPT-2e beats the 50-member ECMWF ENS mean on RMSE and CRPS with 30 members. For intraday trading that requires high-frequency updates, EPT-2 RR refreshes up to 24 times per day, compared with the two-to-four daily runs available from traditional NWP.

The Jua for Energy platform also runs Aurora, GraphCast, ECMWF HRES, ECMWF ENS, ECMWF AIFS, and 19 other models under a single schema. Traders can benchmark any combination on their own region and variable in under 5 minutes.

How does EPT-2 differ from Microsoft Aurora and Google DeepMind GraphCast?

Three structural differences separate EPT-2 from Aurora and GraphCast. First, EPT-2 uses native any-Δt forecasting and predicts at arbitrary lead times without rolling forward in fixed 6-hour increments, so it avoids the error compounding that affects Aurora and GraphCast. Second, EPT-2e is a productised 30-member ensemble that beats the 50-member ECMWF ENS mean on RMSE and CRPS, while neither Aurora nor GraphCast ships a productised ensemble equivalent. Third, Aurora produces no surface solar radiation output, and EPT-2 covers SSRD natively.

At the platform level, Aurora and GraphCast are research outputs from large AI labs. Jua for Energy operates as a production platform with a 24-runs-per-day operational refresh, Athena’s natural-language agent layer, and a 25-model benchmarking surface on which both models run as guests.

Can I integrate Jua forecasts into my own trading models and pipelines?

Integration with existing trading infrastructure is straightforward. The Jua for Energy REST API and Python SDK (pip install jua) expose all 25+ models through a single schema with Apache Arrow support for large payloads. Hindcast data is available across multiple Jua and third-party models for backtesting.

Quant teams pipe forecasts directly into systematic strategies, and utilities and trading houses route them into existing dispatch and risk tools. An ENTSO-E integration provides European grid data in the same workspace. Integration work that takes a quarter to build from raw research-model outputs typically stands up in days on the Jua for Energy platform.

Conclusion and Next Steps for Evaluation

The 2026 benchmark record shows a clear pattern. EPT-2 leads all evaluated AI weather forecasting models on RMSE across 10 m wind, 100 m wind, 2 m temperature, and SSRD at every lead time from 0 to 240 hours. EPT-2e leads on probabilistic skill, beating the 50-member ECMWF ENS mean on CRPS with 30 members. The physics-constrained architecture that produces these results, a general spatiotemporal transformer foundation model that learns conservation laws from observational data, is the same architecture Jua is applying to other physical-economy domains beyond atmospheric prediction.

For energy traders, meteorologists, and quant developers, the practical implication is a single platform where EPT-2, EPT-2e, Aurora, GraphCast, ECMWF HRES, ECMWF ENS, and 19 additional models run under one schema. All models are benchmarked transparently, refreshed up to 24 times per day, and accessible via a natural-language agent that resolves queries in 90 seconds. The numbers speak for themselves. Run them on your own region and variables at athena.jua.ai.

Request a live benchmark on your region and compare EPT-2 head-to-head with your current forecast provider.

Want to talk to the team
behind the writing?

Book a demo to see EPT-2 and Athena in production, or read the open papers behind the work.