Predicting Hailstorms Through Machine Learning Approach Using Multiple Source Data Analysis
摘要
There were 39 hailstorms in the regions of DKI Jakarta, Banten, and West Java from 2010 to 2019. The high frequency of these occurrences underscores the need to study hailstorms as part of extreme weather phenomena, especially considering their potential impact on agriculture, infrastructure, and public safety. Hail events are intimately linked to cloud microphysics and growth processes that can be effectively analyzed using weather radar, Himawari-8 satellite, and ERA-5. This research, conducted from 2015 to 2022, explores an innovative approach by employing a random forest classifier to distinguish three classes of events: no rain, rain, and hail. The input data quality for the prediction model was found to be relatively robust, leading to congruent input and output data. The radar predictor’s significant influence on outcomes was revealed, displaying a higher maximum CSI value on CMAX predictor by 0.33 for event differentiation with threshold 44 dBZ. Although the Himawari-8 predictor was significant for predicting events, its uniform data distribution posed challenges in event differentiation. The combination of multiple source data, coupled with the machine learning approach, makes it possible to greatly improve the robustness of hailstorms prediction compared to any single data commonly used in operational forecasting. Also, that multiple source data produced the most accurate rather than other combinations with accuracy, recall, precision, and F1 scores of 0.9084, 0.9084, 0.9125, and 0.9047, respectively. The accuracy achieved by a random forest model brings encouraging prospects for future research.