Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways for Wind and Power Traders
- EPT-2 beats ECMWF HRES at every lead time for 10 m and 100 m wind, scored directly against more than 10,000 ground stations with no post-processing.
- EPT-2e outperforms the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time, strengthening probabilistic control of imbalance risk.
- Native any-Δt forecasting and up to 24 daily updates from EPT2-RR keep signals fresher than traditional NWP, which typically runs 2–4 times per day.
- A four-percentage-point accuracy gain on 100 m wind delivers about €1.5 M annual P&L impact for a 1 GW wind portfolio in European markets.
- Book a demo with Jua to benchmark EPT-2 against your current provider and quantify the upside for your own assets.
StationBench: How EPT-2 Is Tested on Real Wind Sites
EPT-2 performance on 10 m and 100 m wind uses the StationBench methodology, which scores model outputs directly against observations from more than 10,000 real ground stations, with no post-processing or station fine-tuning. StationBench differs from grid-point interpolation scores because it tests the model where the physics matters operationally, at the surface and at hub height.
Across the full 0–240 hour lead-time range, EPT-2 outperforms ECMWF HRES at every lead time for both 10 m and 100 m wind on RMSE. EPT-2e, the 10-member ensemble variant, also beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time, which directly supports imbalance-cost management through better probabilistic skill.
Two operational differentiators compound this accuracy advantage. EPT-2 uses native any-Δt forecasting and is trained to predict at arbitrary lead times rather than rolling forward in fixed 6-hour increments. This structure avoids the stepwise error accumulation that affects Aurora and most AI peers, which roll in 6-hour steps and compound error at each roll. That architectural edge then supports higher update frequency: EPT-2 runs four times per day, while EPT2-RR, Jua’s rapid-refresh variant, updates up to 24 times per day, compared with 2–4 runs per day for traditional NWP. Traders relying on traditional runs often work with forecasts that are 6–12 hours old, while EPT2-RR keeps the signal current throughout the trading day.
Test EPT-2’s StationBench performance on your own wind sites and variables. Compare it against 25+ models in under 5 minutes at athena.jua.ai.
Model Rankings: EPT-2 vs HRES, ENS, and Aurora on Wind
To put EPT-2’s performance in context, the table below compares it with leading industry models using StationBench results from arXiv 2507.09703. All RMSE and CRPS figures are station-verified with no post-processing.
| Model | 10 m Wind RMSE rank (0–240 h) | 100 m Wind RMSE rank (0–240 h) | Probabilistic skill (CRPS) vs. ENS mean |
|---|---|---|---|
| EPT-2 | Outperforms HRES at every lead time | Outperforms HRES at every lead time | — |
| EPT-2e | Outperforms ECMWF ENS mean at virtually every lead time | Outperforms ECMWF ENS mean at virtually every lead time | Beats 50-member ENS mean on CRPS at virtually every lead time |
| ECMWF HRES | 40-year NWP benchmark | 40-year NWP benchmark | — |
| Microsoft Aurora | Loses to EPT-2 across full 0–240 h range | Loses to EPT-2 across full 0–240 h range | No productised ensemble |
EPT-2e’s advantage over the 50-member ECMWF ENS mean matters directly for trading desks. The ENS has long served as the gold standard for probabilistic NWP, so a 10-member AI ensemble that surpasses it on RMSE and CRPS at almost every lead time signals a structural shift in probabilistic wind forecasting. Aurora currently offers no productised ensemble and no surface solar radiation output, which creates a material gap for portfolios that combine wind and solar exposure.
Explore how these rankings translate to your books. Use Athena to compare EPT-2 against HRES, ENS, Aurora, and more than 20 additional models at athena.jua.ai.
Financial Impact: Error Reduction on 100 m Wind
A four-percentage-point reduction in 100 m wind forecast RMSE has a direct cash impact, not just a cleaner benchmark chart. For a 1 GW wind portfolio in European day-ahead and intraday markets, that accuracy gain delivers about €1.5 M in annual P&L impact under typical hedging and imbalance-penalty structures. The table below links accuracy gains to estimated yearly savings.
| Portfolio size | Accuracy gain (percentage points) | Estimated annual saving |
|---|---|---|
| 1 GW wind | 4 pp | ~€1.5 M/year |
| 1 GW solar | 4 pp | ~€3 M/year |
The mechanism is straightforward. European balancing markets assess imbalance penalties on the gap between nominated generation and actual delivery. A more accurate 100 m wind forecast at 48-hour lead time shrinks the nomination error submitted to the day-ahead market. A more accurate intraday update, delivered hours before the next traditional NWP run, lets the desk revise nominations before the imbalance locks in. Both effects cut the volume exposed to settlement prices, which can reach multiples of the day-ahead clearing price during high volatility.
Probabilistic forecasting, measured by CRPS, adds a second layer of savings. CRPS (Continuous Ranked Probability Score) captures the full distributional accuracy of an ensemble forecast, including the mean, spread, and tails. A tighter, better-calibrated P10–P90 interval on 100 m wind lowers the capital tied up in tail-risk hedging. EPT-2e’s stronger CRPS scores relative to the ECMWF ENS mean translate into reduced hedge volumes and lower option premiums for balancing-responsible parties.
Quantify the P&L impact for your portfolio. Book a demo to see EPT-2’s accuracy advantage on your own wind region.
Probabilistic Forecasting and Imbalance Penalties
Imbalance penalties depend on nomination error and settlement price. Accurate point forecasts reduce nomination error, while well-calibrated probabilistic forecasts reduce the cost of the residual error that remains. CRPS measures this probabilistic calibration, and EPT-2e improves both sides of the equation.
EPT-2e beats the 50-member ECMWF ENS mean on CRPS at virtually every lead time, evaluated on StationBench against a large global station network. For a balancing-responsible party, a tighter ensemble spread at 24-hour lead time narrows and improves the P10–P90 interval on expected wind generation. That improvement reduces the balancing energy volume that must be pre-contracted as insurance against tail outcomes.
Intraday update cadence amplifies this benefit. EPT-2e refreshes four times per day, while traditional NWP typically delivers 2–4 global runs in 24 hours. Between those runs, the desk trades on a forecast that can be 6–12 hours old. EPT2-RR, the rapid-refresh deterministic variant, updates up to 24 times per day and delivers revised wind signals hours before the next traditional NWP cycle completes. In intraday markets, where trade windows last only a few hours and model revisions move prices, that cadence advantage becomes a direct trading edge.
Athena, Jua’s AI agent, converts raw EPT-2 physics predictions into trading-relevant analysis by reading market context and surfacing forecast revisions at the moment they matter. A correction alert fires when EPT-2 revises its own output between runs. A divergence alert fires when EPT-2 and ECMWF ENS disagree on a key variable. Both arrive as notifications, so the trade window opens with a clear signal instead of a missed move.
See how this probabilistic edge behaves on your assets. Run your own comparisons against 25+ models in minutes at athena.jua.ai.
Frequently Asked Questions
What is the published RMSE advantage of EPT-2 over ECMWF HRES on 100 m wind?
The published results in arXiv 2507.09703 document EPT-2’s RMSE advantage across the full 0–240 hour range, validated against more than 10,000 ground stations using the open-source StationBench methodology. The evaluation applies no post-processing or station fine-tuning, so the model is scored directly against surface observations at hub-relevant heights. EPT-2e, the ensemble variant, also improves RMSE and CRPS relative to the 50-member ECMWF ENS mean, which links directly to imbalance-penalty reduction for balancing-responsible parties.
How many model runs per day does EPT2-RR deliver compared with traditional NWP?
EPT2-RR, Jua’s rapid-refresh deterministic variant, delivers up to 24 model runs per day, and EPT-2e updates four times per day. Traditional NWP systems, including ECMWF HRES, usually deliver 2–4 global runs per 24 hours because numerical weather prediction is compute-intensive. A single NWP simulation consumes about 8,400 kWh and costs €1,000–€20,000 on HPC infrastructure. A single EPT-2 inference runs on one GPU in minutes at roughly 0.25 kWh and $0.20–$15, which is about four orders of magnitude cheaper at run time. That cost asymmetry makes 24 daily updates operationally viable and gives Jua for Energy customers revised wind signals well before the next traditional NWP cycle completes.
Can Athena convert a natural-language query into a live wind-trading backtest in under two minutes?
Athena, Jua’s AI agent instrumented with the Jua for Energy tool surface, resolves typical natural-language queries in about 90 seconds and completes backtests in about 5 minutes. A trader can ask, for example, “backtest a wind-ramp strategy on EPT-2e over the last two winters in northern Germany” and receive a full backtest report without writing code. The same agent generates benchmarks, briefings, and custom widgets on request. Quant developers who prefer programmatic access can use the Python SDK, pip install jua, and run their own backtests directly against years of historical EPT-2 and third-party model forecasts. The 25-model benchmarking surface, including ECMWF HRES, ECMWF ENS, Aurora, and GraphCast, is available on the same platform at athena.jua.ai.
Conclusion: From Benchmark to Trade on a Single Platform
EPT-2’s RMSE and CRPS results on 10 m and 100 m wind, validated on StationBench and published in arXiv 2507.09703, already power live production workflows through Jua for Energy. These are not lab-only scores; they are the forecasts that desks use today.
Jua is a foundation model and agent company. EPT is a general physics foundation model, and Athena is an AI agent. Jua for Energy is the first applied product built on both and delivers a transparent 25-model benchmarking surface, a 24×/day refresh cadence via EPT2-RR, and an AI agent that turns natural-language questions into live backtests in about 90 seconds, without displacing the ECMWF feeds customers already run.
The four-percentage-point accuracy gain documented earlier translates to about €1.5 M in annual savings for a 1 GW portfolio, with multi-GW portfolios scaling those economics linearly. The economics follow directly from the physics.
Book a demo to see EPT-2 benchmarked head-to-head against your current forecast provider on your own wind region in under 5 minutes.
