An Ensemble Machine Learning Framework for Predicting Agricultural Droughts in the Ganges Delta of Bangladesh
摘要
Agricultural drought poses a significant threat to food security in many agriculture-dependent countries worldwide. Bangladesh, being one of them, relies heavily on agriculture and frequently experiences crop production deficits due to drought. This study developed a high-resolution, season-specific ensemble machine learning framework to predict agricultural drought severity across three primary cropping seasons—Kharif-1, Kharif-2, and Rabi—in the inactive Ganges Deltaic region. Meherpur District was selected as the case study due to its geomorphological isolation and heavy reliance on groundwater, which together amplify its vulnerability. Using twenty one agrometeorological and socioeconomic parameters, along with the Vegetation Health Index (VHI)—which integrates the Vegetation Condition Index (VCI) and the Temperature Condition Index (TCI)—this study employed five optimized regression models—Random Forest (RF), Extra Trees, AdaBoost, Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM)—within an ensemble framework to predict drought scenarios for the three cropping seasons. Although the model performance yielded moderate predictive accuracy, with R² values ranging from 0.53 to 0.55, the drought prediction revealed significant spatiotemporal variability albeit in different intensity. Extreme drought areas expanded from 6.27% in Kharif-1 to 33.89% in Kharif-2 and persisted at 21.28% during Rabi, with seasonal drought gradually shifting from south to north. Unions such as Meherpur, Amjhupi, and Pirojpur consistently remained high-risk drought zones across all seasons. This research provides a practical and policy-relevant framework for drought preparedness, climate-resilient agriculture, and localized drought early warning in moribund deltaic regions and their surrounding areas.
Graphical AbstractThis study explores an innovative workflow to predict drought conditions in the inactive Ganges delta using Machine Learning(ML) models. This research utilizes a wide range of geospatial data, incorporating 21 agrometeorological and socioeconomic parameters along with the Vegetation Health Index (VHI), derived from the Vegetation Condition Index (VCI) and the Temperature Condition Index (TCI). Data processing involves all of the variables for a consistent 100 m resolution and training machine learning to achieve the best performance metrics to predicting drought conditions in major three cropping seasons. The logical flow provides explanation from data source to the ensemble machine learning framework, incorporating Random Forest, Extra Trees, AdaBoost, XGBoost, and LightGBM models. The high-resolution seasonal drought maps illustrate the spatial and temporal dynamics of agricultural drought in Meherpur District during Kharif-1, Kharif-2, and Rabi seasons. Key results show that extreme drought severity expanded from 6% in Kharif-1 to 34% in Kharif-2, with spatial shifts in drought hotspots identified through Moran’s I analysis. The graphical abstract underscores the study’s central finding: ensemble ML methods provide accurate, fine-scale, season-specific drought forecasts that can be adopted to use the same strategies in different drought-prone regions. This research also provides a data driven results interpretation to support this approach for predicting drought conditions.