Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways for European Wind Traders
- EPT-2 beats ECMWF HRES on 100m wind speed accuracy across every lead time from 0–240 hours, verified against more than 10,000 European ground stations on StationBench with no post-processing applied.
- EPT-2e, the ensemble variant, outperforms the 50-member ECMWF ENS mean on CRPS at virtually every lead time and delivers tighter probabilistic calibration for energy trading decisions.
- The accuracy edge holds for both offshore and onshore sites, with the largest gains in complex terrain where EPT-2’s physics-constrained architecture captures orographic effects more reliably than traditional NWP.
- Statistical post-processing cuts forecast error by 10–39 percent on top of EPT-2’s baseline, which converts into substantial annual P&L savings for multi-GW wind portfolios under European imbalance and hedging structures.
- See EPT-2 on your own sites by running live 100m wind benchmarks on your European region and comparing EPT-2 to your current provider in under five minutes.
Latest StationBench Results: 100m Wind Accuracy by Lead Time
The tables below show RMSE and CRPS rankings for 100m wind speed across Europe, drawn from arXiv:2507.09703, the EPT-2 technical report. All figures are verified against more than 10,000 ground stations on open-source StationBench. Lower values indicate better performance.
0–6 Hour Lead Time (Very Short Range)
| Model | RMSE Rank | CRPS Rank | Notes |
|---|---|---|---|
| EPT-2 | Leading | — | Deterministic, outperforms ECMWF HRES across the horizon |
| EPT-2e | — | Leading | Ensemble mean, outperforms ECMWF ENS mean on CRPS |
| ECMWF HRES | 2nd | — | 40-year NWP benchmark, deterministic |
| ECMWF ENS | — | 2nd | 50-member ensemble, probabilistic NWP reference |
| ICON-EU | 3rd | 3rd | DWD regional European model |
| Microsoft Aurora | 4th | — | No productised ensemble, fixed 6-hour roll-forward compounds error |
6–24 Hour Lead Time (Short Range)
| Model | RMSE Rank | CRPS Rank | Notes |
|---|---|---|---|
| EPT-2 | Leading | — | Native any-Δt, does not roll forward in fixed steps |
| EPT-2e | — | Leading | Outperforms ECMWF ENS mean on CRPS |
| ECMWF HRES | 2nd | — | Universal benchmark |
| ECMWF ENS | — | 2nd | 50-member probabilistic reference |
| ICON-EU | 3rd | 3rd | Higher-resolution European regional model |
| Microsoft Aurora | 4th | — | Loses to EPT-2 on 100m wind across the full forecast horizon |
24–72 Hour Lead Time (Medium Range)
| Model | RMSE Rank | CRPS Rank | Notes |
|---|---|---|---|
| EPT-2 | Leading | — | Outperforms ECMWF HRES across lead times |
| EPT-2e | — | Leading | Ensemble skill exceeds ECMWF ENS mean on CRPS |
| ECMWF HRES | 2nd | — | Strongest NWP deterministic competitor |
| ECMWF ENS | — | 2nd | 50-member ensemble |
| ICON-EU | 3rd | 3rd | Skill degrades faster than global models at longer ranges |
| Microsoft Aurora | 4th | — | No SSRD output, 100m wind skill trails EPT-2 across the horizon |
72–240 Hour Lead Time (Extended Range)
| Model | RMSE Rank | CRPS Rank | Notes |
|---|---|---|---|
| EPT-2 | Leading | — | Maintains lead over ECMWF HRES out to 240 h |
| EPT-2e | — | Leading | 60-day ensemble horizon, outperforms ENS mean at virtually every lead time |
| ECMWF HRES | 2nd | — | 10-day deterministic horizon |
| ECMWF ENS | — | 2nd | 15-day ensemble horizon |
| ICON-EU | 3rd | 3rd | Regional model, limited extended-range skill |
| Microsoft Aurora | 4th | — | Research output, no productised operational schedule |
These results use StationBench verification against more than 10,000 stations, with no post-processing applied to any model. RMSE (root mean square error) measures deterministic point accuracy, and CRPS (continuous ranked probability score) measures probabilistic calibration. Lower values are better on both metrics.
Offshore and Onshore Performance Across European Wind Regimes
Offshore and onshore 100m wind sites create different forecasting challenges. Offshore locations such as the North Sea, Baltic Sea, and Irish Sea have lower surface roughness and more homogeneous boundary-layer dynamics. That structure usually produces lower absolute RMSE values across all models. Onshore sites in complex terrain, including the Alps, Pyrenees, Scandinavian highlands, and Iberian plateau, introduce orographic channeling, thermal inversions, and wake effects. These factors amplify model error, especially at lead times beyond 24 hours.
EPT-2’s StationBench evaluation spans both regimes, with the station network covering coastal, offshore-adjacent, and complex-terrain sites across the European domain. The skill advantage EPT-2 holds over ECMWF HRES appears in both regimes, and the absolute error gap widens in complex terrain. EPT-2’s physics-constrained latent representation captures orographic forcing more accurately than fixed-grid NWP at equivalent resolution.
Jua for Energy delivers forecasts at 1 km native resolution via EPT-2 HRRR over Europe. For site-specific applications such as individual wind farm assessments and hub-height profiling at 11 levels from 10 m to 200 m, the platform supports downscaling to 1 km resolution. That scale matches operational needs for offshore cluster planning and complex-terrain dispatch decisions.
Post-Processing Calibration: 10–39% Error Reduction on Top of EPT-2
Raw 100m wind output is only the starting point for trading decisions. Statistical post-processing, including bias correction, ensemble calibration, and model output statistics (MOS), reduces RMSE and CRPS relative to uncalibrated output. Across European wind sites, post-processing delivers 10–39 percent error reductions, depending on site characteristics, lead time, and the baseline model.
The trading economics follow directly from these gains. Under typical European hedging and imbalance penalty structures, a 1 GW wind portfolio with higher forecast accuracy saves substantial costs each year. When post-processing reaches the upper end of the 10–39 percent improvement range on a multi-GW portfolio, the cumulative annual P&L impact scales quickly. Jua’s forecasts therefore carry an estimated material P&L impact per gigawatt annually in European energy markets.
EPT-2’s physics-constrained architecture also reduces the systematic biases that post-processing must correct. EPT-2 learns conservation laws such as mass, momentum, and energy directly from observational data instead of approximating them through grid-cell discretization. The remaining error structure is less correlated across lead times. Calibration then focuses on random error instead of systematic model drift, which makes post-processing more effective.
Quantify the calibration uplift by scheduling a live benchmark on your own European region.
From Benchmark to Trading Desk: How Jua for Energy Uses EPT-2
Jua operates as a foundation model and agent company, and Jua for Energy is the first applied product. It runs on EPT, a general spatiotemporal transformer foundation model for physical systems, and Athena, an AI agent instrumented with the energy-trader tool surface. The relationship mirrors Anthropic and Claude Code, pairing a horizontal platform with a flagship vertical product.
Inside Jua for Energy, the StationBench results above function as live, reproducible benchmarks that run on any region a customer selects. The benchmarking surface places more than 25 models on a single platform, including EPT-2, EPT-2e, ECMWF HRES, ECMWF ENS, ICON-EU, and Microsoft Aurora, and returns a head-to-head accuracy comparison in seconds. Meteorologists who start sceptical of vendor accuracy claims often become internal champions once they run the benchmark on their highest-stakes region.
Athena, the AI agent, converts a natural-language request such as “show the 100m wind forecast spread across models for the German Bight tonight” into a briefing, a benchmark, or a full backtest in about 90 seconds. Backtests against years of historical forecasts complete in roughly five minutes. EPT-2 HRRR provides frequent updates over Europe at high native resolution, compared to less frequent daily runs from traditional NWP. EPT-2e updates multiple times per day. Divergence alerts trigger the moment two models disagree on 100m wind in a subscribed zone, surfacing a trade window before the market re-prices.
Customers including Axpo, TotalEnergies, Statkraft, EnBW, EDF, and Hydro-Québec already execute daily trading decisions on the platform. The Python SDK installs via pip install jua, and the REST API exposes all 25+ models through a single schema with Apache Arrow support for large payloads.
Frequently Asked Questions
Which model leads 100m wind accuracy at 24–72 hour lead times in Europe?
EPT-2 leads on RMSE and EPT-2e leads on CRPS across this medium-range window in Europe, verified against the StationBench ground-station network with no post-processing applied. ECMWF HRES is the strongest NWP competitor in this range and ranks second on deterministic RMSE. ICON-EU ranks third. Microsoft Aurora trails EPT-2 on 100m wind across the full forecast horizon. The 24–72 hour window is commercially critical for day-ahead and multi-day power market positioning, and the EPT-2 advantage here translates directly into reduced imbalance costs for wind portfolio operators.
How much do post-processing techniques improve 100m wind forecasts?
The 10–39 percent error reductions mentioned earlier depend on site type, terrain complexity, and lead time. Offshore sites with homogeneous boundary-layer dynamics tend to see smaller absolute gains because the raw model error is already lower. Complex-terrain onshore sites such as Alpine valleys, Pyrenean passes, and Scandinavian fjords see larger calibration gains because systematic orographic biases are more pronounced and more correctable. EPT-2’s physics-constrained architecture reduces the systematic bias component that post-processing must address, which makes calibration more effective on top of an already more accurate baseline.
What resolution do Jua models deliver for European wind sites?
EPT-2 HRRR delivers forecasts at 1 km native resolution over Europe, covering wind at 11 height levels from 10 m to 200 m, the full range of commercial wind turbine hub heights. For site-specific applications such as individual wind farm assessments or offshore cluster planning, Jua for Energy supports downscaling to 1 km resolution. This compares to 9 km for ECMWF HRES and approximately 25 km for Microsoft Aurora at published resolution. Wind at 100 m is a native output variable, not interpolated from adjacent levels, which improves hub-height accuracy at modern turbine installations.
Where can traders run their own 100m wind benchmark in under five minutes?
The live benchmarking surface inside Jua for Energy puts more than 25 models, including EPT-2, EPT-2e, ECMWF HRES, ECMWF ENS, ICON-EU, and Microsoft Aurora, on a single platform. A trader or meteorologist selects a European region, selects 100m wind as the variable, and the platform returns a head-to-head RMSE and CRPS comparison against ground-station observations in seconds. Athena can extend this into a full multi-year backtest in about five minutes. The benchmark is available to any prospect during evaluation and to all customers post-procurement for ongoing model surveillance. Booking a demo remains the fastest path to running the benchmark on your own portfolio region.
Conclusion: Turn StationBench Accuracy into Trading P&L
EPT-2 outperforms ECMWF HRES on 100m wind at every lead time from 0 to 240 hours, verified against StationBench’s European ground stations. EPT-2e beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. Post-processing calibration then adds a further 10–39 percent error reduction on top of that baseline. For a 1 GW wind portfolio, this accuracy uplift converts into substantial annual savings in recovered hedging and imbalance costs.
Generic model descriptions and vendor graphics no longer suffice for procurement decisions or daily trading operations. Station-verified numbers at operationally relevant lead times now exist, remain reproducible, and run on any European region in under five minutes inside Jua for Energy.
Run your region-specific accuracy test by booking a demo to compare 25+ models head-to-head and see your 100m wind numbers before the next trade window opens.
