Comparative assessment of classical and machine learning approaches for rainfall data restoration
摘要
Incorporating a comprehensive long-term hydrological data is a crucial aspect of conducting water resource management studies. This approach enhances the precision of hydrological models. This article aims to investigate and compare various classical and machine learning (ML) methods for recovering missing rainfall data. The study focuses on five mountainous basins in the Central Alborz Ranges in Iran, utilizing 30 years of data. The classical methods used in the study include arithmetic average (AA), linear regression (LR), multiple linear regression (MLR), inverse distance weighting (IDW), kriging with three different semi-variogram and normal ratio (NR) models, and a suggested linear regression-arithmetic average (LR-AA) method. The ultimate goal is to identify suitable methods for accurately recovering missing rainfall data in the studied region. Several machine learning methods were employed to restore precipitation data, such as artificial neural networks (ANN), support vector regression (SVR), M5 trees, and, as a novel approach, two types of adaptive neuro-fuzzy inference systems (ANFIS). To ensure that the selected duration does not have any potential impact, three intervals of artificial gaps have been incorporated to minimize the uncertainties in recovery period. These periods include 1990–1993, 2002–2005, and 2011–2014. In addition, a Social Choice method was coupled with the evaluation criteria to enhance the comparison process. In general, the results indicate that machine learning methods outperform than the classical approaches. For example, during the gap of 2002–2005 in the Karaj basin, the SVR method is the most effective method with RMSE, NSE and