Towards More Reliable Gridded Precipitation Estimates: Gauge-Based Multi-Scale Evaluation and Machine Learning Bias Correction
摘要
Precipitation is a key component in water resource management. However, in many regions, ground-based observations remain limited. Consequently, gridded satellite precipitation datasets provide an effective alternative. This study focuses on two main.(i) to evaluate and compare the performance of four high-resolution precipitation products—PERSIANN-CDR, CHIRPS, ERA5-Land, and GPM IMERG-Final (GPM-F)—against observations from 12 rain gauges in the Moulouya Basin (Morocco); and (ii) to enhance the accuracy of the best-performing product through machine-learning-based bias correction. The evaluation covers daily, monthly, seasonal, and annual time scales, and both pixel- and basin-level spatial domains, applying standard hydrometeorological metrics (CC, R², RBias, RMSE, NRMSE, NSE). GPM-F consistently shows superior performance, particularly at the daily scale. To further improve its accuracy, three machine learning models (MLs), including Artificial Neural Network (ANN), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost), are applied for bias correction. Results show that model performance is scale-dependent. At the daily scale, RF performs best, increasing the median correlation coefficient (CC) to ~ 0.8, raising the coefficient of determination (R²) from 0.24 to 0.65, and reducing the root mean square error (RMSE) by 61.8%. At the monthly scale, XGBoost slightly outperforms RF. For seasonal data, XGBoost achieves the highest accuracy, with CC up to 0.96, R² above 0.85, RMSE reduced by 80.4%, and consistently high Nash–Sutcliffe efficiency (NSE) values (> 0.85). These findings demonstrate the potential of ML-based correction to substantially reduce bias and error, providing more reliable precipitation inputs for hydrological applications and aligning with the growing use of AI-driven approaches in environmental sciences.
Graphical AbstractGraphical Abstract Description: The graphical abstract illustrates the methodological framework designed to evaluate and enhance satellite-based precipitation data for hydrological applications over the Moulouya Basin, Morocco. The process begins with a multi-scale spatio-temporal analysis of precipitation datasets derived from three sources: reanalysis data, satellite-based products (PERSIANN-CDR, CHIRPS, ERA5, and GPM-F), and ground-based observations. This initial phase involves data harmonization, quality control, and statistical comparison across different spatial and temporal scales to identify the most reliable satellite product. Based on this evaluation, GPM-F is selected as the most accurate dataset, particularly at the daily scale. To further improve its precision, three machine learning (ML) models: Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Artificial Neural Network (ANN), are applied for bias correction. Each model learns the relationship between satellite and ground observations to adjust the systematic bias and enhance predictive performance. The final output consists of bias-corrected precipitation time series produced at daily, monthly, and seasonal scales. These corrected datasets exhibit substantial improvements in correlation, determination coefficients, and error reduction metrics, with performance varying by scale. The refined precipitation products are suitable for hydrological modeling, flood risk assessment, and water resource management, demonstrating how ML-based correction effectively transforms biased satellite estimates into reliable, high-accuracy datasets for operational use.