Written by: Olivier Lam, Physical AI Team, Jua.ai AG
Key Takeaways for European Energy Desks
- Consumer weather sites provide city-level data that is too coarse, stale, and opaque for professional energy-trading decisions across Europe’s varied climate zones.
- Physics-based foundation models like EPT-2 embed conservation laws directly in their architecture and deliver physically consistent 2 m temperature forecasts that outperform ECMWF HRES and Microsoft Aurora at every lead time.
- EPT-2 RR updates up to 24 times per day at up to 5 km resolution, closing the six-hour staleness gaps that exist in traditional NWP cycles and consumer-site refreshes.
- Benchmarks against more than 10,000 ground stations show EPT-2e beats the 50-member ECMWF ENS mean on both RMSE and CRPS, with all results reproducible on the Jua platform without post-processing.
- Energy teams can validate forecasts in seconds and run backtests in minutes via Athena; book a demo with Jua to see EPT-2 head-to-head against your current provider.
Why Physics-Based Foundation Models Fit Energy Trading
Physics-based foundation models form a distinct class of AI system tailored to real-world dynamics. They differ from large language models because an ideal Earth foundation model must incorporate physical consistency, including conservation laws, symmetry, and causality. This design avoids predictions driven by non-physical relationships or spurious correlations.
Standard transformers applied directly to atmospheric data can output fields that break mass, momentum, and energy conservation. A physics-constrained foundation model avoids this failure mode because the conservation laws live inside the representation rather than as a patch applied afterward.
Jua operates as a foundation model and agent company. Its Earth Physics Transformer (EPT) family is a general spatiotemporal transformer foundation model that learns the governing physics of complex systems directly from observational data. The architecture remains domain-agnostic, while the data and fine-tune define each physical system. Atmospheric prediction is the first system EPT has been fine-tuned for, and Jua for Energy is the first applied product built on EPT and on Athena, Jua’s AI agent.
Athena plans, reasons, and calls tools to turn a natural-language question into a benchmark or briefing in about 90 seconds and a backtest in about 5 minutes, or into a custom widget. This pairing of a physics-constrained foundation model with an agent-assisted workflow separates Jua for Energy from both consumer weather sites and raw AI-weather research outputs that force customers to build their own ingestion pipeline.
Existing physics-based weather and climate models rely on imperfect parameterizations of high-resolution microphysics, including clouds and aerosols, ocean eddies, and land–atmosphere interactions. EPT addresses this limitation by learning these dynamics directly from observational data instead of depending on hand-coded parameterization schemes.
How EPT-2 Generates Hourly 2 m Temperature Forecasts
EPT-2, the deterministic flagship model in the EPT family, is trained on more than 5 petabytes of weather and climate data from over 120 distinct sources. These sources include geostationary and polar-orbiting satellites, surface station networks, national radar networks, ocean buoys, ERA5 reanalysis, and operational ECMWF initial-condition fields. The model learns a latent representation of atmospheric state and then integrates that state forward in time faster than the atmosphere itself evolves.
A typical EPT-2 workflow for a European energy desk follows a clear pattern. At each new model cycle, EPT-2 ingests the latest observational state of the atmosphere. EPT2-HRRR then produces a deterministic 2 m temperature forecast at up to 5 km native resolution across Europe, at arbitrary lead times from the current hour out to 20 days. The model does this without rolling forward in fixed 6-hour increments like Aurora and most peer AI models. Rolling introduces error at every step, and EPT-2’s native any-Δt architecture avoids that accumulation entirely.
EPT-2 RR, the rapid-refresh variant, updates up to 24 times per day. EPT-2 HRRR delivers the same high-cadence output at up to 5 km resolution over Europe. For a trader watching the German intraday market, this cadence means a fresh 2 m temperature signal every hour instead of every six hours and not only at the 00z and 12z cycles that define the traditional NWP schedule.
The output is accessible through three integration paths that match different workflows. The Jua platform’s REST API supports programmatic access. The pip install jua package on PyPI supports Python-native development. Athena’s natural-language interface supports non-technical users who still need direct access to the forecasts. This flexibility allows a quant developer to pipe EPT-2 2 m temperature forecasts directly into an internal risk model in the same afternoon they first evaluate the SDK.
Evidence from EPT-2 Temperature Benchmarks (arXiv 2507.09703)
EPT-2’s performance on 2 m temperature appears in the peer-reviewed technical report arXiv:2507.09703. The evaluation methodology uses open-source StationBench and compares forecasts against more than 10,000 real ground stations, with no post-processing or station fine-tuning applied to EPT-2’s outputs.
EPT-2 outperforms ECMWF HRES on 2 m temperature across all lead times from 0 to 240 hours. EPT-2 also outperforms Microsoft Aurora on 2 m temperature across the full 0–240 hour range. EPT-2e, the ensemble variant, beats the 50-member ECMWF ENS mean on both RMSE (root mean square error) and CRPS (continuous ranked probability score) at virtually every lead time.
These results do not come from vendor-produced graphics. Prospects can reproduce them on the Jua platform’s live benchmarking surface by selecting their own region, variable, and time window and then viewing a head-to-head comparison in seconds.
Why Traditional NWP Creates Stale Temperature Signals
ECMWF’s IFS and AIFS operational models run four global forecast cycles per day at 00, 06, 12, and 18 UTC, with 2 m temperature included in each cycle’s surface fields. For the 00z and 12z IFS cycles, 2 m temperature forecasts appear at hourly steps from 0 to 90 hours, then 3-hourly steps to 144 hours, then 6-hourly steps to 360 hours. The 06z and 18z cycles provide 3-hourly steps only to 144 hours.
Four cycles per day create a maximum inter-run gap of six hours. In practice, the energy industry receives roughly four global forecasts in any 24-hour period. Between runs, traders stare at stale numbers. A cold front accelerating across the North Sea at 14:00 UTC will not appear in a revised ECMWF 2 m temperature signal until the 18z run completes and disseminates, which can arrive hours later.
Consumer weather sites inherit this staleness because they consume NWP output instead of producing independent forecasts. When ECMWF updates, consumer sites eventually reflect the change. When ECMWF sits between runs, consumer sites display the last available output, often without indicating its age or the size of the revision since the previous cycle.
EPT-2 RR updates up to 24 times per day. A trader running Jua for Energy alongside an ECMWF subscription sees the next temperature revision hours before the next traditional NWP run lands. Correction alerts fire the moment EPT-2 revises its own output between cycles, which signals that the atmosphere has moved while the market has not yet repriced.
Why Current Forecast Workflows Slow Energy Desks
The standard energy-desk approach to hourly temperature forecasts across Europe functions as a manual assembly line. A meteorologist or analyst downloads raw grib files from ECMWF’s MARS archive, processes them through an in-house pipeline, cross-references a second model such as GFS, ICON, or a subscribed AI output, and then stitches the result into a spreadsheet or terminal screen. When a model updates mid-day, the desk often notices only because someone else has already traded on it.
Consumer weather sites do not fix this workflow. They add a parallel data stream with no API, no ensemble, no model-comparison surface, and no audit trail. A meteorologist who must defend a temperature forecast to a risk committee cannot credibly cite a consumer site. The data is not benchmarked, not versioned, and not traceable to a physical model with documented skill scores.
Jua for Energy consolidates more than 25 models on a single platform with a unified schema and a single API. The platform includes 10 proprietary AI models from the EPT family and 15 third-party NWP and AI models such as ECMWF HRES, ECMWF ENS, ECMWF AIFS, NOAA GFS, DWD ICON, Microsoft Aurora, and GFS GraphCast. A meteorologist who wants to evaluate model disagreement on 2 m temperature over Iberia during a heat-advection event can run a head-to-head benchmark across all models in seconds on their own region and variable without writing code.
Athena, Jua’s AI agent instrumented with the Jua for Energy tool surface, converts a natural-language question into a benchmark or briefing in about 90 seconds and a backtest in about 5 minutes. Trading houses and quant desks describe Athena as “another headcount, for free.”
Book a demo to run a live benchmark on your own region and variable.
Why Frequent Updates Break Traditional Compute Budgets
Traditional NWP runs only four times per day because of hard economic constraints, not design preference. A single NWP simulation consumes approximately 8,400 kWh of compute and costs €1,000–€20,000 to run on high-performance computing infrastructure. The European supercomputer that runs ECMWF’s full algorithm can sustain two complete global runs per day, and supplementary runs raise the total to four. No utility or trading house can afford to run NWP more frequently on its own infrastructure.
A single EPT-2 inference runs on a single GPU in minutes at approximately 0.25 kWh and $0.20–$15 per simulation. EPT-2 was trained on 8 × H100 GPUs over 10 days, while Microsoft Aurora required 32 × A100 GPUs over 18 days. EPT-2 therefore used four times fewer GPUs and a substantially shorter training cycle. The inference cost differential reaches roughly four orders of magnitude. That cost profile makes 24 updates per day economically viable for EPT-2 RR, while traditional NWP cannot reach that cadence.
EPT-2 vs ECMWF HRES vs Consumer Sites: Key Facts
The table below compares EPT-2 against ECMWF HRES on metrics that matter for European energy trading. All EPT-2 benchmark figures come from arXiv:2507.09703. ECMWF operational specifications come from ECMWF open-data documentation.
| Capability | EPT-2 (Jua for Energy) | ECMWF HRES | Consumer Weather Sites |
|---|---|---|---|
| 2 m temperature accuracy (0–240 h, RMSE vs. 10,000+ ground stations) | Outperforms ECMWF HRES at every lead time | 40-year benchmark; gold standard for NWP | Not benchmarked; no published skill scores |
| Update frequency | up to 24×/day (EPT-2 RR) | 4×/day (00, 06, 12, 18 UTC) | Downstream of NWP; update timing opaque |
| Native spatial resolution (Europe) | Up to 5 km (EPT-2 HRRR) | 9 km (HRES) | City-level; no published grid specification |
| Probabilistic / ensemble output | EPT-2e beats 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time | ENS: 50-member gold standard for probabilistic NWP | None |
Risk Checks for Hourly 2 m Temperature Providers
Any energy-trading team that evaluates hourly 2 m temperature forecast providers should apply the following criteria before creating a workflow dependency.
Benchmark transparency. The provider must publish skill scores against independent ground-truth observations, not only vendor-produced graphics. Acceptable methodologies include StationBench or equivalent open-source evaluation against station networks of at least several thousand sites, with no post-processing applied to the model output.
Update cadence documentation. The provider must state the number of forecast cycles per day, the dissemination latency after each cycle completes, and the inter-run gap during which the forecast remains stale. Four cycles per day with a six-hour inter-run gap form the NWP baseline. Any higher cadence should come with documentation of the underlying model architecture that enables it.
Physics grounding. AI models that do not embed conservation laws for mass, momentum, and energy can produce physically inconsistent outputs. An ideal Earth foundation model must incorporate physical consistency to avoid predictions based on non-physical relationships or spurious correlations. Teams should require providers to document how physical constraints are enforced at the representation level rather than as post-processing.
API and hindcast access. For defensible energy-trading use, the provider must expose forecast data through a documented API with schema stability and must supply hindcast data for backtesting across at least two years of historical forecast cycles.
Peer-reviewed validation. Accuracy claims must be traceable to a published technical report. For EPT-2, the reference remains arXiv:2507.09703.
Frequently Asked Questions
What is the most accurate hourly weather forecast for Europe?
For professional energy-trading applications, accuracy depends on RMSE and CRPS against independent ground-truth station networks rather than consumer ratings. EPT-2, the deterministic flagship model in Jua’s EPT family, already demonstrates a documented accuracy advantage over ECMWF HRES on 2 m temperature across the full 0–240 hour range. EPT-2e, the ensemble variant, improves on the 50-member ECMWF ENS mean on both RMSE and CRPS at virtually every lead time. Consumer sites do not publish comparable benchmarks.
Why do consumer weather sites fail for energy trading?
Consumer weather sites target public guidance instead of defensible trading inputs. They deliver city-level temperature data without published skill scores, ensemble spreads, or API access. Their update schedule follows NWP cycles, typically four times per day, and they rarely flag the age of the displayed forecast or the size of the revision since the previous cycle. They cannot integrate into systematic trading models, do not supply hindcast data for backtesting, and cannot be benchmarked transparently against alternative models.
How does Jua for Energy differ from subscribing directly to ECMWF?
Jua for Energy does not replace an ECMWF subscription, and most serious customers run both. ECMWF HRES and ENS, along with ECMWF AIFS, run natively on the Jua platform alongside EPT-2 and 22 other models under a unified schema and a single API. Jua for Energy replaces the plumbing around the ECMWF feed, including the in-house grib pipeline, manual benchmarking, morning-briefing assembly, and dashboard stitching. EPT-2 RR also delivers hourly updates, providing fresh 2 m temperature signals in the hours between ECMWF’s four daily cycles.
What is Jua, and is it a weather company?
Jua operates as a foundation model and agent company rather than a traditional weather provider. The Earth Physics Transformer (EPT) family functions as a general physics foundation model with a domain-agnostic architecture, and Athena serves as the AI agent. Jua for Energy is the first applied product built on both components. Atmospheric prediction is the first physical system EPT has been fine-tuned for, and energy trading is the first market Athena supports.
How quickly can I validate EPT-2 against my current forecast provider?
The live benchmark on the Jua platform returns a head-to-head accuracy comparison in seconds. A prospect selects their region, their variable, such as 2 m temperature over the Iberian Peninsula, and their current provider. The platform then returns the comparison immediately. Backtests against years of historical forecast cycles run in about 5 minutes via Athena. This benchmark usually acts as the deal trigger because the numbers speak before a procurement conversation begins.
Conclusion: How to Choose Defensible Hourly Temperature Intelligence
Consumer weather sites do not provide a viable source of hourly 2 m temperature data for European energy trading. They lack benchmark transparency, clear update cadence documentation, ensemble output, API access, and physics grounding. The professional alternative uses a physics-constrained foundation model with documented skill scores, high-cadence updates, and a productised workflow layer.
EPT-2 delivers this alternative. It offers documented accuracy advantages over ECMWF HRES at every lead time, supports high-frequency updates through the rapid-refresh cycle described above, and appears in arXiv:2507.09703 with validation against a large ground-station network. A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves about €1.5 M per year.
Teams can run benchmarks on their own regions and variables on the Jua platform. View forecasts head-to-head against more than 25 models at athena.jua.ai, or book a demo to see EPT-2 against your current provider in a guided session.
