<p>Heavy precipitation events, while often associated with devastating floods and societal disruption, represent a critical water resource if managed effectively. Accurate early detection and differentiation of such events in short-term forecasts remain a significant challenge, particularly in urban areas. This study defines heavy precipitation using thresholds where daily precipitation exceeds the 99th, 95th, and 90th percentiles at one, two, and three stations, respectively, with at least half of the stations recording precipitation. To ensure distinct events, days within three days before and after heavy precipitation were excluded, and non-heavy precipitation days were randomly selected, resulting in a dataset of 184 days (92 heavy and 92 non-heavy precipitation days). The dataset was divided into training (75%) and testing (25%) sets, with atmospheric data analyzed from one to five days prior to events. Feature selection methods and averaging techniques were compared to optimize model performance. The Random Forest (RF) model, using input data from 1 to 4 days prior, combined with the "Both" averaging method and Chi-Square feature selection, demonstrated superior performance, achieving an accuracy of 0.848, recall of 0.880, and precision of 0.846. Logistic regression also performed competitively, achieving an accuracy of 0.804. Key predictors included zonal wind and specific humidity, with Outgoing Longwave Radiation (OLR) and vapor flux anomalies revealing a consistent sequence: negative OLR anomalies indicated strong initial convection, followed by intensification and eastward movement of a Mediterranean Sea cyclone, enhancing vapor flux and leading to heavy precipitation. These findings offer actionable insights for improving forecasting and early warning systems in regions prone to extreme weather.</p> Graphical Abstract <p>This graphical abstract presents an artistic and visually engaging representation of a machine learning approach to analyze heavy precipitation events in southwest Iran. The illustration creatively integrates scientific concepts with vivid imagery—featuring a robot intently studying a book titled "Machine Learning", symbolizing the application of advanced algorithms in weather prediction. Binary code manifests as precipitation, symbolizing the machine learning classification method. Above the study area, elegant machine learning icons blend seamlessly with swirling and raining clouds, emphasizing the fusion of atmospheric science and artificial intelligence. The abstract highlights key findings, with the Random Forest (RF) model emerging as the most accurate (0.848 accuracy) when using atmospheric data from 1–4 days before events, outperforming Logistic Regression (0.804 accuracy) when using atmospheric data from 1–2 days before events. This comparison is playfully depicted by a thinking figure in the top-right corner, pondering RF’s superiority. Critical predictors like zonal wind and specific humidity are visually emphasized, while spatial analyses of outgoing longwave radiation (OLR) and vapor flux reveal dynamic pre-storm patterns through a time-delay sequence. The backdrop showcases a map of study area, with rain clouds over southwest Iran and an umbrella shielding two children symbolizing the study’s focus on extreme precipitation. The central hypothesis is boldly overlaid on the region, tying the artistic elements to the research’s core objective: improving forecasting through machine learning. This visually rich abstract not only communicates complex methodologies but also makes the science accessible, blending creativity with rigorous analysis to underscore the study’s innovation and practical significance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting Heavy Precipitation in Southwest Iran: A Machine Learning Classification Approach with Atmospheric Precursors and Feature Optimization

  • Kokab Shahgholian,
  • Javad Bazrafshan,
  • Parviz Irannejad,
  • Dariush Ranazadeh,
  • Vijay P. Singh

摘要

Heavy precipitation events, while often associated with devastating floods and societal disruption, represent a critical water resource if managed effectively. Accurate early detection and differentiation of such events in short-term forecasts remain a significant challenge, particularly in urban areas. This study defines heavy precipitation using thresholds where daily precipitation exceeds the 99th, 95th, and 90th percentiles at one, two, and three stations, respectively, with at least half of the stations recording precipitation. To ensure distinct events, days within three days before and after heavy precipitation were excluded, and non-heavy precipitation days were randomly selected, resulting in a dataset of 184 days (92 heavy and 92 non-heavy precipitation days). The dataset was divided into training (75%) and testing (25%) sets, with atmospheric data analyzed from one to five days prior to events. Feature selection methods and averaging techniques were compared to optimize model performance. The Random Forest (RF) model, using input data from 1 to 4 days prior, combined with the "Both" averaging method and Chi-Square feature selection, demonstrated superior performance, achieving an accuracy of 0.848, recall of 0.880, and precision of 0.846. Logistic regression also performed competitively, achieving an accuracy of 0.804. Key predictors included zonal wind and specific humidity, with Outgoing Longwave Radiation (OLR) and vapor flux anomalies revealing a consistent sequence: negative OLR anomalies indicated strong initial convection, followed by intensification and eastward movement of a Mediterranean Sea cyclone, enhancing vapor flux and leading to heavy precipitation. These findings offer actionable insights for improving forecasting and early warning systems in regions prone to extreme weather.

Graphical Abstract

This graphical abstract presents an artistic and visually engaging representation of a machine learning approach to analyze heavy precipitation events in southwest Iran. The illustration creatively integrates scientific concepts with vivid imagery—featuring a robot intently studying a book titled "Machine Learning", symbolizing the application of advanced algorithms in weather prediction. Binary code manifests as precipitation, symbolizing the machine learning classification method. Above the study area, elegant machine learning icons blend seamlessly with swirling and raining clouds, emphasizing the fusion of atmospheric science and artificial intelligence. The abstract highlights key findings, with the Random Forest (RF) model emerging as the most accurate (0.848 accuracy) when using atmospheric data from 1–4 days before events, outperforming Logistic Regression (0.804 accuracy) when using atmospheric data from 1–2 days before events. This comparison is playfully depicted by a thinking figure in the top-right corner, pondering RF’s superiority. Critical predictors like zonal wind and specific humidity are visually emphasized, while spatial analyses of outgoing longwave radiation (OLR) and vapor flux reveal dynamic pre-storm patterns through a time-delay sequence. The backdrop showcases a map of study area, with rain clouds over southwest Iran and an umbrella shielding two children symbolizing the study’s focus on extreme precipitation. The central hypothesis is boldly overlaid on the region, tying the artistic elements to the research’s core objective: improving forecasting through machine learning. This visually rich abstract not only communicates complex methodologies but also makes the science accessible, blending creativity with rigorous analysis to underscore the study’s innovation and practical significance.