Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways
- Verification of AI atmospheric models against European ground stations now sits at the core of regulatory and commercial risk management for utilities and trading houses.
- StationBench delivers an open benchmark framework that compares raw model output with real SYNOP and METAR observations across more than 10,000 stations, without post-processing.
- EPT-2 and EPT-1.5 beat ECMWF HRES and other leading AI models on energy-relevant variables under the same transparent verification protocol.
- Published StationBench results supply the reproducible evidence that EU AI Act conformity assessments require for high-risk AI systems in energy infrastructure.
- Utilities and traders can run live head-to-head benchmarks on any European region in under five minutes — schedule a demo with Jua to see results for your portfolio.
How European Ground-Station Networks Support AI Verification
European observational networks provide the ground truth that AI atmospheric models must match. Each network offers different station density, variable coverage, and access rules, so any serious verification approach needs to navigate this patchwork.
The four principal frameworks — ESA Φ-lab, EUMETSAT, ACTRIS, and EUROCONTROL — illustrate how fragmented this landscape is for AI model verification.
ESA Φ-lab coordinates satellite–ground validation across ESA member states and focuses on Earth observation data fusion rather than direct NWP verification. Station access remains project-specific and not uniformly public. EUMETSAT maintains ground-truth infrastructure for its satellite products, including SYNOP-linked surface validation sites across Europe, with data available to registered users under the EUMETSAT Data Policy. ACTRIS runs European supersites with atmospheric profiling for aerosol and cloud research, and its data is publicly accessible through the ACTRIS Data Centre. EUROCONTROL aggregates METAR observations from civil aviation stations, providing surface wind, temperature, visibility, and pressure through EUROCONTROL NEST and B2B services.
National SYNOP networks run by DWD, Météo-France, the UK Met Office, KNMI, and their peers add several thousand more stations. These data flow under WMO agreements and are accessible through national portals and the ECMWF MARS archive. Together, these networks form the observational backbone that any credible AI model verification methodology must use.
To see how this backbone behaves for your assets, run a live benchmark for your own European region and variables.
Core Verification Metrics Used in Europe
Four metrics dominate European AI model verification practice, and each one answers a different question about forecast quality. Root Mean Square Error (RMSE) measures the average magnitude of forecast error in physical units, such as metres per second for wind or degrees Celsius for temperature. It penalises large errors quadratically, so it highlights extreme misses.
Mean Absolute Error (MAE) measures the average absolute deviation without that quadratic penalty. It complements RMSE by showing typical error size rather than focusing on extremes. Bias captures systematic over- or under-prediction across the verification period. Energy applications care about bias because a persistent directional error in wind or solar forecasts turns directly into imbalance costs.
Continuous Ranked Probability Score (CRPS) extends the idea of RMSE to probabilistic forecasts. It measures the integrated difference between the forecast cumulative distribution function and the observed outcome. Lower CRPS means the ensemble spread is both accurate and well calibrated.
Rigorous European workflows apply these metrics directly to raw station observations such as SYNOP reports, METAR records, or proprietary feeds. They do not adjust model output to match local station climatology first. Post-processing can hide structural model weaknesses and inflate skill scores, which makes results hard to transfer to new regions or variables.
The StationBench methodology, described next, enforces this rule explicitly. It evaluates raw model output against raw observations, with no station fine-tuning. This standard gives regulators and risk committees something they can audit, and it is the standard used to evaluate EPT-2. That auditability requirement is now a legal obligation under the EU AI Act.
To see these metrics on your own portfolio, compare EPT-2 against ECMWF HRES for your region.
EU AI Act Requirements for Weather and Energy Models
The EU AI Act, in force since August 2024 with phased obligations through 2026 and 2027, classifies AI systems used in critical infrastructure as high-risk under Annex III. Energy systems fall squarely into this category. High-risk AI systems must pass conformity assessments that demand detailed technical documentation, including training data provenance, model architecture, and performance validation methods.
These systems must also support ongoing monitoring and logging of outputs so operators can trace behaviour over time. Transparency obligations give competent authorities the right to audit model behaviour. Human oversight requirements ensure that operators can step in when outputs look unreliable, instead of relying blindly on automated decisions.
For AI atmospheric models used in dispatch, balancing-responsible-party (BRP) obligations, or trading, these rules create a direct need for ground-station verification evidence. Accuracy claims that rely only on vendor graphics or grid-point comparisons against reanalysis data fall short of the auditability standard. Ground-station verification against a named, reproducible dataset, with published RMSE, CRPS, and bias figures, provides the documentary trail that a conformity assessment expects. EPT-2’s StationBench results, published on arXiv, supply exactly this type of evidence.
If you need to document compliance for your own stack, generate auditable verification reports for your regions.
StationBench Methodology and EPT-2 Results
The regulatory need for auditable verification evidence shapes the design of StationBench. StationBench is an open-source benchmark framework that covers more than 10,000 weather stations worldwide, with optional regional analysis for Europe, for AI atmospheric model evaluation. It evaluates raw model output, interpolated to station coordinates but otherwise unchanged, against raw SYNOP and METAR observations. It then computes RMSE, CRPS, MAE, and bias across variables and lead times.
The framework is fully reproducible. The station list, observation preprocessing code, and evaluation scripts are public, so any team can rerun or extend the benchmark independently. EPT-2, documented in arXiv:2507.09703, outperforms ECMWF HRES on energy-relevant variables across the 0–240 hour horizon. EPT-2e, the ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. EPT-1.5, documented in arXiv:2410.15076, outperforms GraphCast, FuXi, Pangu-Weather, and ECMWF HRES on European wind and temperature under the same no-post-processing protocol.
StationBench is the only framework among the main European verification systems that combines open public access with published AI-specific benchmarks, a combination that matters for regulatory compliance. The table below compares the five principal European verification frameworks across four dimensions relevant to AI model evaluation.
| Framework | Station Count | Variables | Public Access | AI-Specific Benchmarks |
|---|---|---|---|---|
| ESA Φ-lab | Project-specific, no fixed public count | EO-derived surface and atmospheric variables | Project-restricted | None published for NWP or AI weather models |
| EUMETSAT | ~2,000+ SYNOP-linked validation sites across Europe | Surface temperature, wind, humidity, radiation | Registered users under EUMETSAT Data Policy | Satellite product validation only, no AI NWP benchmarks |
| ACTRIS | Supersites across Europe | Aerosol, cloud, trace gases, atmospheric profiling | Public via ACTRIS Data Centre | None for AI weather prediction models |
| EUROCONTROL | Civil aviation METAR stations | Wind, temperature, visibility, pressure, cloud ceiling | Available via EUROCONTROL B2B services | None published for AI atmospheric models |
| StationBench | 10,000+ worldwide (with optional regional analysis for Europe) | 2 m temperature and 10 m wind speed | Open-source, scripts and station list publicly available | EPT-2 vs ECMWF HRES and others with published metrics |
To explore these benchmarks for your own assets, run a StationBench-based comparison in under five minutes.
Practical Workflow for Utilities and Trading Houses
A meteorologist or quant developer at a European utility or trading house can run a live head-to-head benchmark on any European region and variable in under five minutes on the Jua platform at athena.jua.ai. The workflow is straightforward. You select a region such as the German onshore wind corridor or the Iberian solar belt, choose a variable like 100 m wind speed or surface solar radiation, and select EPT-2 alongside ECMWF HRES and any third-party AI model. The platform then returns a head-to-head RMSE and CRPS comparison on the spot, without grib file processing, in-house pipelines, or delays.
EPT-2e, the ensemble variant, updates four times per day and provides probabilistic forecast distributions that map directly to BRP imbalance-cost calculations and options-desk hedging workflows. Those distributions gain extra value from Jua’s ~5 km resolution over Europe (EPT2-HRRR), which resolves mesoscale wind and solar variability that 9 km ECMWF HRES grids smooth out. Because the 25+ models on the Jua platform, including ECMWF HRES, ECMWF ENS, ECMWF AIFS, Microsoft Aurora, GFS GraphCast, DWD ICON, and the full EPT family, run under a unified schema, teams can switch or compare models without re-engineering pipelines. Jua already serves major utilities across four continents, including some of Europe’s largest energy companies, along with commodity traders and hedge funds.
Start your benchmark now and see model comparisons for your portfolio in minutes.
FAQ
What makes StationBench different from standard NWP verification frameworks?
Most NWP verification frameworks apply bias correction or statistical post-processing to model output before computing skill scores against station observations. StationBench does not. Raw EPT-2 output is interpolated to station coordinates and evaluated directly against raw SYNOP and METAR records across the StationBench network. This constraint keeps results fully reproducible for any third party and prevents inflation of scores through local climatological adjustments. The open-source scripts and station list are publicly available, which helps satisfy the auditability requirement that EU AI Act conformity assessments impose on high-risk AI systems in energy infrastructure.
Does EPT-2 outperform ECMWF HRES on all lead times and variables relevant to energy trading?
Yes. As documented in the StationBench results, EPT-2 consistently outperforms ECMWF HRES across all lead times on the variables that matter for energy trading, including wind speed, temperature, and solar radiation. EPT-2e, the ensemble variant, also improves probabilistic skill relative to traditional ensembles. Microsoft Aurora has no surface solar radiation output, so EPT-2 remains the only AI model with a complete verification record on key variables against European ground stations. These findings appear in the EPT-2 and EPT-1.5 technical reports on arXiv.
How does physics-constrained architecture reduce the risk of outputs that violate conservation laws?
Standard transformers applied directly to atmospheric data can produce outputs that break mass, momentum, and energy conservation, which the real atmosphere never does. EPT instead learns the governing physics of complex systems from observational data in a latent representation that evolves forward in time. This architecture enforces conservation laws by design, so the model cannot generate physically impossible states even when it produces novel patterns. The ground-station verification record on StationBench then shows that this physics-constrained design translates into real-world forecast accuracy, not just theoretical consistency.
Conclusion
StationBench provides an open verification path for AI atmospheric models against its global station network, with optional regional analysis for Europe. EPT-2 and EPT-1.5 have published results on that benchmark, meeting both the technical standard that meteorologists expect and the auditability obligation that the EU AI Act places on high-risk AI systems in energy infrastructure. To see how this applies to your own assets, run the benchmark for your regions and variables at athena.jua.ai.
