Weather Forecasting

Atmospheric Model Skill Scores: HSS, ACC & BSS Guide

Olivier Lam·May 21, 2026
Atmospheric Model Skill Scores: HSS, ACC & BSS Guide 2026

Written by: Olivier Lam, Physical AI Team, Jua.ai AG | Last updated: June 28, 2026

Key Takeaways

  • Atmospheric model skill scores (ACC, BSS, HSS, RPSS) quantify forecast improvement over climatology and support direct comparison of NWP and AI models on energy-relevant variables.
  • ACC above 0.6 on 500 hPa height marks the operational threshold for skillful medium-range forecasts that support wind-ramp and temperature-spread trading decisions.
  • BSS above 0.1 and HSS above 0.3 provide actionable probabilistic and categorical guidance for day-ahead hedging and automated alert systems in renewable portfolios.
  • RPSS above 0.1 on temperature terciles at 7–10 days enables systematic multi-category positioning for gas-demand and renewable-generation strategies.
  • Benchmark EPT-2e against 25+ models in a live demo on your own region and variables in under five minutes.

Skill Scores at a Glance for Energy Trading

ScoreFormulaCommon ThresholdEnergy-Trading Interpretation
ACC(Σ f′o′) / √(Σ f′² · Σ o′²), where f′ and o′ are forecast and observed anomalies from climatology>0.6 considered skillful for medium-range, >0.8 operationally reliableACC >0.6 at 240 h on 500 hPa height indicates a model resolves synoptic patterns that drive wind-ramp and temperature-spread events. Each 0.1-point gain above 0.6 corresponds to measurable imbalance-cost reduction on a 1 GW portfolio.
BSS1 − (BS / BS_clim), where BS = (1/N)Σ(p_i − o_i)²>0.0 beats climatology, >0.1 operationally useful, >0.25 strong probabilistic skillBSS >0.1 on 10 m wind exceedance events supports tighter bid-offer positioning in day-ahead power markets. BSS >0.25 supports ensemble-based hedging strategies on renewable generation.
HSS(PC − E) / (1 − E), where PC is proportion correct and E is expected proportion correct by chance>0.0 beats random, >0.3 useful for categorical events, >0.5 strongHSS >0.3 on wind-ramp categorical events (for example >10 m/s threshold) reduces missed-ramp penalties. HSS >0.5 supports automated alert thresholds in trading systems.
RPSS1 − (RPS / RPS_clim), where RPS = Σ(CDF_forecast − CDF_observed)²>0.0 beats climatology, >0.1 useful multi-category skill, >0.2 strong ensemble discriminationRPSS >0.1 on temperature tercile forecasts at 7–10 days supports gas-demand positioning. RPSS >0.2 enables systematic multi-category renewable-generation strategies.

Each of these skill scores plays a distinct role in forecast evaluation. The next sections walk through them in detail, starting with ACC, the foundational metric for synoptic-scale verification.

Anomaly Correlation Coefficient (ACC)

ACC measures the linear correlation between forecast and observed anomalies, both computed relative to the climatological mean. The formula is ACC = (Σ f′o′) / √(Σ f′² · Σ o′²). Values range from −1 to 1, where 1.0 is a perfect forecast and 0.0 indicates no skill beyond climatology.

The 500 hPa geopotential height field is the standard ACC benchmark variable in medium-range NWP verification. It captures synoptic-scale dynamics that propagate into surface wind and temperature forecasts. A threshold of 0.6 has historically marked the boundary of useful medium-range prediction.

ACC Thresholds for Tradeable Medium-Range Skill

ACC above 0.6 represents the operational threshold for skillful medium-range forecasting established by the NWP community over four decades. Below 0.6, synoptic-pattern errors grow large enough that surface-variable forecasts derived from the 500 hPa field carry substantial uncertainty. Above 0.6, the large-scale flow is sufficiently resolved so wind-ramp and temperature-spread events can be anticipated with tradeable confidence. ACC above 0.8 is considered operationally reliable for day-ahead and multi-day energy positioning and extends the window in which energy traders can act on model guidance.

Brier Skill Score (BSS)

The Brier Score (BS) measures the mean squared error of probabilistic forecasts for binary events: BS = (1/N)Σ(p_i − o_i)², where p_i is the forecast probability and o_i is the binary outcome. The Brier Skill Score normalises this against a climatological reference: BSS = 1 − (BS / BS_clim). A BSS of 0.0 matches climatology, 1.0 is perfect, and negative values indicate the model performs worse than climatology.

Operational BSS Levels for Wind and Solar

In operational NWP verification, BSS >0.0 confirms the model beats climatology on the event in question. BSS >0.1 marks the point where probabilistic guidance becomes operationally useful, meaning the ensemble’s probability estimates are sufficiently reliable to inform hedging or dispatch decisions. BSS >0.25 represents strong probabilistic skill and supports systematic ensemble-based trading strategies. EPT-2e, Jua’s ensemble variant, beats the 50-member ECMWF ENS mean on RMSE and CRPS at virtually every lead time, which corresponds to BSS performance above the operationally useful threshold across the variables that drive renewable-generation forecasts.

How to Read BSS for Energy Applications

BSS is event-specific. A model can score BSS 0.3 on a wind-exceedance event and BSS 0.05 on a precipitation-threshold event. Interpretation requires the variable, the threshold, the lead time, and the climatological reference period. For energy-trading applications, the most relevant BSS evaluations cover 10 m and 100 m wind exceedance (wind-ramp risk), surface solar radiation below a generation threshold (solar dip risk), and 2 m temperature extremes (gas-demand spikes). A BSS improvement of 0.05 on a 1 GW wind portfolio’s ramp-event forecast corresponds to a material reduction in imbalance penalties under European balancing-responsible-party structures.

Heidke Skill Score (HSS)

HSS evaluates categorical forecasts, meaning events classified into discrete bins, against the proportion correct expected by chance. The formula is HSS = (PC − E) / (1 − E), where PC is the proportion of correct categorical forecasts and E is the expected proportion correct under random forecasting. HSS ranges from −∞ to 1, where 0.0 equals chance and 1.0 is perfect.

For energy trading, HSS applies to categorical wind-ramp events such as wind speed crossing a turbine cut-in or cut-out threshold, temperature tercile classification (below-normal, near-normal, above-normal), and solar-irradiance regime classification. HSS >0.3 on wind-ramp categorical events is the threshold at which automated alert systems can be calibrated without excessive false-positive rates. HSS >0.5 supports rule-based trading strategies that trigger on categorical model output. EPT-2’s outperformance of ECMWF HRES on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation across 0–240 hour lead times translates directly into HSS gains on the categorical events that matter most to wind and solar portfolio managers.

Ranked Probability Skill Score (RPSS)

RPSS extends the Brier Score to multi-category probabilistic forecasts by comparing cumulative distribution functions. The Ranked Probability Score is RPS = Σ(CDF_forecast − CDF_observed)², and RPSS = 1 − (RPS / RPS_clim). RPSS >0.0 beats climatology, >0.1 is operationally useful for multi-category ensemble guidance, and >0.2 indicates strong ensemble discrimination across the full probability distribution.

RPSS is the preferred metric for ensemble forecast evaluation in energy markets because it penalises both overconfident and underconfident probability distributions. A model that assigns 90% probability to the wrong tercile scores worse on RPSS than one that assigns 40% to the correct tercile. For gas-demand positioning at 7–10 day lead times, RPSS >0.1 on temperature tercile forecasts marks the point where ensemble guidance adds value over climatological priors. EPT-2e’s 30-member ensemble beats the 50-member ECMWF ENS mean on CRPS at virtually every lead time, with CRPS as the continuous generalisation of RPS, confirming RPSS performance above the operationally useful threshold.

Run your own head-to-head skill-score comparison on your region and variables against 25+ models on the Jua platform in under five minutes.

2026 EPT-2e versus ECMWF ENS Benchmark

MetricEPT-2e (Jua, 2026)ECMWF ENS (50-member mean)
RMSE vs. ENS mean (10 m wind, 0–240 h)Beats ENS mean at virtually every lead timeReference benchmark
CRPS vs. ENS mean (all key variables, 0–240 h)Beats ENS mean at virtually every lead timeReference benchmark
Ensemble members30 members50 members
Native spatial resolutionUp to 5 km (EPT-2 HRRR, Europe)9 km (ENS)
Update frequency4×/day2–4×/day
Inference cost per simulation~$0.20–$15, ~0.25 kWh, single GPU~€1,000–€20,000, ~8,400 kWh, HPC cluster

As documented in the performance data above, EPT-2e consistently outperforms the ENS mean on RMSE and CRPS while using fewer ensemble members. EPT-2 HRRR forecasts natively at up to 5 km resolution over Europe and resolves mesoscale features that 9 km NWP grids smooth over. The cost asymmetry reaches four orders of magnitude: a single EPT-2 inference runs at approximately 0.25 kWh and $0.20–$15 on a single GPU, compared with about 8,400 kWh and €1,000–€20,000 for an equivalent NWP simulation on HPC infrastructure.

ECMWF, GFS, and AI Models on Energy Variables

ECMWF HRES has led deterministic NWP for four decades and remains the universal benchmark. NOAA GFS is the free deterministic baseline used across the energy industry. Both run their full algorithm two to four times per day. The relevant 2026 comparison focuses on how ECMWF and GFS perform relative to AI models on the variables that drive energy-market P&L.

EPT-2 outperforms ECMWF HRES on every lead time from 0 to 240 hours on 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation, the four variables with the largest direct impact on wind and solar portfolio P&L. EPT-1.5 outperforms GraphCast, FuXi, Pangu-Weather, and ECMWF HRES on European wind and temperature. EPT-2 also beats Microsoft Aurora on 10 m wind, 100 m wind, and 2 m temperature across the full 0–240 hour range, while Aurora has no surface solar radiation output. Jua for Energy runs alongside ECMWF for serious customers, who keep their ECMWF subscription. Jua displaces the plumbing around the incumbent feed: the in-house grib pipeline, the manual benchmarking, and the morning-briefing analyst.

Mapping Skill Scores to 1 GW Portfolio P&L

Skill-Score ThresholdVariable / EventAnnual P&L Impact (1 GW portfolio)
ACC >0.6 sustained to 240 h500 hPa height → synoptic wind patternEnables multi-day wind positioning. A 4 percentage-point accuracy gain on a 1 GW wind portfolio saves about €1.5 M/year.
BSS >0.1 on wind exceedance10 m / 100 m wind ramp eventsSupports tighter imbalance-cost exposure. The same 4 percentage-point accuracy gain on a 1 GW solar portfolio saves about €3 M/year.
HSS >0.3 on wind-ramp categoricalTurbine cut-in / cut-out threshold crossingsReduces missed-ramp penalties under BRP structures and supports automated alert calibration.
RPSS >0.1 on temperature tercile (7–10 day)Gas-demand tercile classificationEnables systematic multi-day gas positioning. RPSS >0.2 supports rule-based ensemble strategies at scale.

A 1 GW wind portfolio that gains four percentage points of forecast accuracy saves approximately €1.5 M per year under typical European hedging and imbalance-penalty structures. A 1 GW solar portfolio at the same accuracy gain saves approximately €3 M per year. Customers operating multi-GW portfolios scale these economics linearly. The skill-score thresholds in the table above define the quantitative conditions under which each saving becomes achievable and act as operational criteria that separate a model worth trading on from one that is not.

Live Benchmarking on the Jua Platform

Jua operates as a foundation model and agent company, with Jua for Energy as its first applied product. This structure mirrors Anthropic’s relationship to Claude Code, a horizontal AI platform expressed through a flagship vertical product. EPT and Athena are domain-agnostic by architecture, while the atmosphere is the first physical system EPT has been fine-tuned for and energy trading is the first market Athena has been instrumented for.

On the Jua platform, more than 25 models, including 10 proprietary AI models from the EPT family and about 15 third-party NWP and AI models such as ECMWF HRES, ECMWF ENS, NOAA GFS, Microsoft Aurora, and GFS GraphCast, run on a single benchmarking surface. A meteorologist evaluating EPT-2e against their current ensemble provider selects a region, a variable, and a time window. The platform then returns a head-to-head RMSE, CRPS, and skill-score comparison in seconds. Athena, Jua’s AI agent, can extend that benchmark into a full backtest in about five minutes or turn a natural-language question into a briefing, a custom widget, or a portfolio-level analysis in about 90 seconds.

This workflow creates the deal trigger for every serious evaluation because the numbers speak directly to the user. Meteorologists who arrive sceptical of vendor accuracy claims often become internal champions once they run the benchmark themselves on their own region and variable.

Run your own head-to-head comparison on the Jua platform and see EPT-2e against your current forecast provider on ACC, BSS, HSS, and RPSS in under five minutes.

Frequently Asked Questions

How do ACC, BSS, HSS, and RPSS differ?

ACC (Anomaly Correlation Coefficient) measures how well a deterministic forecast captures the spatial pattern of anomalies relative to climatology and serves as the standard metric for synoptic-scale NWP verification. BSS (Brier Skill Score) measures the skill of probabilistic forecasts for binary events relative to a climatological baseline. HSS (Heidke Skill Score) evaluates categorical forecasts, meaning events classified into discrete bins, against the proportion correct expected by chance. RPSS (Ranked Probability Skill Score) extends BSS to multi-category probabilistic forecasts and penalises both overconfident and underconfident ensemble distributions. For energy trading, ACC captures synoptic-pattern skill, BSS supports wind-ramp and solar-dip event probabilities, HSS guides categorical alert calibration, and RPSS underpins multi-day ensemble-based positioning.

Which ACC values matter for medium-range energy forecasting?

ACC above 0.6 is the established threshold for skillful medium-range forecasting on 500 hPa geopotential height. Below 0.6, synoptic-pattern errors become large enough to degrade surface-variable forecasts so wind-ramp and temperature-spread events cannot be anticipated with tradeable confidence. ACC above 0.8 is considered operationally reliable for day-ahead and multi-day energy positioning. Sustaining skill above this threshold extends the window in which energy traders can act on model guidance before the market re-prices.

How should I interpret a Brier Skill Score for wind or solar forecasting?

BSS is always event-specific and reference-specific. A BSS of 0.0 means the model matches climatology on the event in question, while negative values mean it performs worse than climatology. For wind and solar energy applications, BSS above 0.1 on exceedance events, such as wind speed crossing a ramp threshold or solar irradiance falling below a generation threshold, marks the minimum level for operationally useful probabilistic guidance. BSS above 0.25 supports systematic ensemble-based hedging strategies. Interpretation requires the variable, the threshold, the lead time, and the climatological reference period, so a BSS reported without these parameters cannot be compared across models or vendors.

How do ECMWF and AI models like EPT-2 compare on energy-relevant variables in 2026?

ECMWF HRES remains the universal NWP benchmark and the reference against which all models are evaluated. In 2026, EPT-2 outperforms ECMWF HRES on every lead time from 0 to 240 hours on the four variables with the largest direct impact on energy-market P&L: 10 m wind, 100 m wind, 2 m temperature, and surface solar radiation. EPT-2e, the ensemble variant, maintains this performance advantage across the full forecast range while using 30 ensemble members. These results appear in Jua’s peer-reviewed technical reports on arXiv. Jua for Energy runs alongside the incumbent ECMWF feed and replaces the manual pipeline built around it.

How quickly can I benchmark EPT-2 against my current forecast provider on the Jua platform?

The live benchmark on the Jua platform returns a head-to-head accuracy comparison in seconds from the first model selection. A full backtest against years of historical forecasts, run via Athena, completes in about five minutes. The benchmark covers more than 25 models, including ECMWF HRES, ECMWF ENS, NOAA GFS, Microsoft Aurora, and GFS GraphCast, on any region, any variable, and any time window the user specifies. No pre-processing, grib files, or in-house pipeline are required.

Conclusion

Atmospheric model skill scores, specifically ACC, BSS, HSS, and RPSS, act as quantitative conditions that separate a forecast worth trading on from one that is not. In 2026, EPT-2e’s RMSE and CRPS advantages over ECMWF ENS translate into a consistent performance edge, and the Jua platform puts all four scores across more than 25 models on any region and variable in front of a meteorologist, quant developer, or trader in seconds. These skill-score improvements translate to the multi-million-euro annual savings detailed in the P&L mapping above.

Jua operates as a foundation model and agent company, with Jua for Energy as its first applied product. The underlying architecture learns physics as a universal capability, which makes the specific domain, in this case atmospheric forecasting, a variable within a broader framework. The benchmark experience is live, transparent, and completes in under five minutes.

See the results on your portfolio variables in a Jua demo and evaluate EPT-2e head-to-head against your current forecast provider.

View the key takeaways as a web story

Want to talk to the team behind the writing?

Book a demo to see EPT-2 and Athena in production, or read the open papers behind the work.