Intercomparison of Machine Learning and Ingredient-Based Approaches for Identifying Hail-Prone Weather Conditions over Russia
摘要
Hail is one of the most dangerous natural phenomena, whose forecasting and diagnosing is difficult due to its small scale and complex formation. Various approaches exist to diagnose the probability of hail based on the information about large-scale atmospheric conditions. This work presents an intercomparison of several such approaches using ERA5 reanalysis data and a hail database for Russian regions. We investigate the effectiveness of three distinct methodologies: a convolutional neural network (CNN), a gradient boosting on trees (CatBoost) model, and a commonly used ingredient-based approach that utilizes the composite WMAXSHEAR index. Interpretability analysis was conducted using SHAP (Shapley additive explanations) and reparameterization techniques, showing consistency with the conditions characteristic of hail formation. A comparative study of the models’ performance shows that the two ML models significantly outperformed the traditional ingredient-based approach in terms of different skill scores. The results could be used for both short-term hail forecasting and long-term severe weather risk assessment, helping to mitigate the impact of large hail.