Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles

Sobash, R. A., Schwartz, C. S., Ahijevych, D.. (2026). Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles. Weather and Forecasting, doi:https://doi.org/10.1175/waf-d-25-0243.1

Title Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles
Genre Article
Author(s) Ryan A. Sobash, Craig S. Schwartz, David Ahijevych
Abstract We assess the skill of daily machine learning (ML)-based probabilistic convective hazard predictions generated from two experimental medium-range (i.e., extending to day 8) convection-allowing ensemble (CAE) systems to determine if such predictions are skillful over the CONUS. CAE forecasts were generated in real time for events in spring 2023 and 2024, and convective hazard predictions were produced using neural networks, trained with CAE output and observed storm reports (i.e., reports of tornadoes, hail, and convectively induced wind gusts). The skill of the two ML-based hazard forecasts was evaluated relative to climatology, CAE-based updraft helicity (UH) forecasts, and ML-based predictions generated from the NOAA operational Global Ensemble Forecast System (GEFS). The CAE ML-based hazard forecasts outperformed both the UH-based forecasts and climatology through day 8. When evaluated against the GEFS ML, CAE-based hazard forecasts were skillful for days 1-4, beyond which differences were small and not statistically significant. All hazard forecasts were more skillful in spring 2024 than in spring 2023, likely due to stronger large-scale forcing that enhanced predictability. The stronger forcing led to reduced benefit of the CAE-based forecasts relative to the GEFS, with benefits extending to days 3-4 in 2024 but extending through day 5 in 2023. Finally, ML interpretability experiments revealed that CAE storm-scale predictors were weighted more than environmental predictors for days 1-2, while the opposite was true for days 3-5. Together, these results document the skill and value of CAEs for ML-based hazard prediction, as well as interannual skill variations due to different weather regimes that can influence skill differences and NWP system intercomparisons.
Publication Title Weather and Forecasting
Publication Date Aug 1, 2026
Publisher's Version of Record https://doi.org/10.1175/waf-d-25-0243.1
OpenSky Citable URL https://n2t.net/ark:/85065/d7v69q6r
OpenSky Listing View on OpenSky
MMM Affiliations PARC

< Back