Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles
Sobash, R. A., Schwartz, C. S., Ahijevych, D.. (2026). Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles. Weather and Forecasting, doi:https://doi.org/10.1175/waf-d-25-0243.1
| Title | Evaluating the skill of medium-range machine learning-based convective hazard forecasts derived from convection-allowing ensembles |
|---|---|
| Genre | Article |
| Author(s) | Ryan A. Sobash, Craig S. Schwartz, David Ahijevych |
| Abstract | We assess the skill of daily machine learning (ML)-based probabilistic convective hazard predictions generated from two experimental medium-range (i.e., extending to day 8) convection-allowing ensemble (CAE) systems to determine if such predictions are skillful over the CONUS. CAE forecasts were generated in real time for events in spring 2023 and 2024, and convective hazard predictions were produced using neural networks, trained with CAE output and observed storm reports (i.e., reports of tornadoes, hail, and convectively induced wind gusts). The skill of the two ML-based hazard forecasts was evaluated relative to climatology, CAE-based updraft helicity (UH) forecasts, and ML-based predictions generated from the NOAA operational Global Ensemble Forecast System (GEFS). The CAE ML-based hazard forecasts outperformed both the UH-based forecasts and climatology through day 8. When evaluated against the GEFS ML, CAE-based hazard forecasts were skillful for days 1-4, beyond which differences were small and not statistically significant. All hazard forecasts were more skillful in spring 2024 than in spring 2023, likely due to stronger large-scale forcing that enhanced predictability. The stronger forcing led to reduced benefit of the CAE-based forecasts relative to the GEFS, with benefits extending to days 3-4 in 2024 but extending through day 5 in 2023. Finally, ML interpretability experiments revealed that CAE storm-scale predictors were weighted more than environmental predictors for days 1-2, while the opposite was true for days 3-5. Together, these results document the skill and value of CAEs for ML-based hazard prediction, as well as interannual skill variations due to different weather regimes that can influence skill differences and NWP system intercomparisons. |
| Publication Title | Weather and Forecasting |
| Publication Date | Aug 1, 2026 |
| Publisher's Version of Record | https://doi.org/10.1175/waf-d-25-0243.1 |
| OpenSky Citable URL | https://n2t.net/ark:/85065/d7v69q6r |
| OpenSky Listing | View on OpenSky |
| MMM Affiliations | PARC |