<p>Accurate rainfall prediction is crucial for effective weather forecasting and climate modeling. This study aims to assess the effectiveness of a novel hybrid Statistical Downscaling-based Hidden Markov Model-Random Forest Model (SD-based HMM-RF) for rainfall prediction in Selangor, Malaysia. It also examines the best imputation methods for handling missing data, selects predictors for statistical downscaling by reducing dimensionality, and addresses uncertainties in zero-bounded rainfall data. The study utilized observed data (predictand) from 33 rainfall stations and atmospheric data (predictor), covering the period from 2008 to 2018. Seven imputation methods were tested: Mean Imputation (MeI), Median Imputation (MI), Expectation-Maximization (EM) Algorithm, Markov Chain Monte Carlo (MCMC), k-Nearest Neighbor (kNN), Non-iterative Partial Least Square (NIPALS), and Random Forest (RF). Principal Component Analysis (PCA) was used to manage high-dimensional data and select predictors, while HMM was applied to address uncertainties in zero-bounded rainfall data. Five hybrid models: Random Forest (SD-based HMM-RF), Support Vector Machine (SD-based HMM-SVM), Decision Tree (SD-based HMM-DT), k-Nearest Neighbors (SD-based HMM-KNN), and Artificial Neural Networks (SD-based HMM-ANN) were evaluated. Performance metrics, including Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Forecast Error (MFE), Nash-Sutcliffe Efficiency (NSE), Kling-Gupta Efficiency (KGE), and Rank Correlation Coefficient (ρ), were used to identify the most accurate rainfall prediction model. MI emerged as the best-performing method, achieving the lowest RMSE (0.146) and MAE (0.021), along with the highest NSE (1.000) and Pearson correlation coefficient, r (1.000) for all stations. PCA identified five principal components (PC1–PC5) with a cumulative variance of up to 93.254%. HMM analysis determined that three hidden states (K = 3) at iteration 2000 were optimal, with the lowest Bayesian Information Criterion (BIC) of 260,018.30. Among the models, SD-based HMM-RF consistently outperformed others, showing superior accuracy and reliability in rainfall prediction, as indicated by its RMSE (2.298), NSE (0.752), MAE (1.900), near-zero MFE (-0.049), KGE (0.526), and ρ (0.898). By improving rainfall prediction accuracy, the novel hybrid model can enhance early warning systems, inform infrastructure planning, and reduce the economic impact of future flooding events, thereby contributing to more resilient urban development in the region.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating a hybrid statistical Downscaling-based HMM-RF model for enhanced rainfall prediction in Selangor

  • Noor Hamizah Mohamad Sani,
  • Shazlyn Milleana Shaharudin,
  • Muhammad Safwan Ibrahim,
  • Upmanu Lall,
  • Mou Leong Tan

摘要

Accurate rainfall prediction is crucial for effective weather forecasting and climate modeling. This study aims to assess the effectiveness of a novel hybrid Statistical Downscaling-based Hidden Markov Model-Random Forest Model (SD-based HMM-RF) for rainfall prediction in Selangor, Malaysia. It also examines the best imputation methods for handling missing data, selects predictors for statistical downscaling by reducing dimensionality, and addresses uncertainties in zero-bounded rainfall data. The study utilized observed data (predictand) from 33 rainfall stations and atmospheric data (predictor), covering the period from 2008 to 2018. Seven imputation methods were tested: Mean Imputation (MeI), Median Imputation (MI), Expectation-Maximization (EM) Algorithm, Markov Chain Monte Carlo (MCMC), k-Nearest Neighbor (kNN), Non-iterative Partial Least Square (NIPALS), and Random Forest (RF). Principal Component Analysis (PCA) was used to manage high-dimensional data and select predictors, while HMM was applied to address uncertainties in zero-bounded rainfall data. Five hybrid models: Random Forest (SD-based HMM-RF), Support Vector Machine (SD-based HMM-SVM), Decision Tree (SD-based HMM-DT), k-Nearest Neighbors (SD-based HMM-KNN), and Artificial Neural Networks (SD-based HMM-ANN) were evaluated. Performance metrics, including Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Forecast Error (MFE), Nash-Sutcliffe Efficiency (NSE), Kling-Gupta Efficiency (KGE), and Rank Correlation Coefficient (ρ), were used to identify the most accurate rainfall prediction model. MI emerged as the best-performing method, achieving the lowest RMSE (0.146) and MAE (0.021), along with the highest NSE (1.000) and Pearson correlation coefficient, r (1.000) for all stations. PCA identified five principal components (PC1–PC5) with a cumulative variance of up to 93.254%. HMM analysis determined that three hidden states (K = 3) at iteration 2000 were optimal, with the lowest Bayesian Information Criterion (BIC) of 260,018.30. Among the models, SD-based HMM-RF consistently outperformed others, showing superior accuracy and reliability in rainfall prediction, as indicated by its RMSE (2.298), NSE (0.752), MAE (1.900), near-zero MFE (-0.049), KGE (0.526), and ρ (0.898). By improving rainfall prediction accuracy, the novel hybrid model can enhance early warning systems, inform infrastructure planning, and reduce the economic impact of future flooding events, thereby contributing to more resilient urban development in the region.