Machine learning-based landslide risk assessment in Doti district, western Nepal
摘要
In the Himalayas, communities face significant landslide threats, which emphasize the critical need for effective, data-driven risk assessment frameworks. This study introduces a machine learning system designed for mapping landslide susceptibility and vulnerability in Doti District, Nepal. We compiled a comprehensive inventory of 1,147 confirmed landslides from 1997 to 2023, along with 13 causative factors that were carefully screened for multicollinearity; all VIF values ranged from 1.04 to 2.94. We evaluated four different models: Logistic Regression, Support Vector Machine, Random Forest, and XGBoost. Among them, XGBoost achieved the highest AUC of 0.96, an accuracy of 88.5%, and a Kappa of 0.77. When we validated the model with 155 recent landslides from 2020 to 2024, it produced an AUC of 0.854, demonstrating strong generalization and minimal overfitting. The feature importance analysis revealed that elevation (0.162), NDVI (0.138), slope (0.118), geology (0.114), and rainfall (0.096) were the key predictors, accounting for nearly 70% of the total importance. Our susceptibility mapping indicated that 18.89% of the district is situated in high to very high-risk zones, particularly along the Main Central Thrust and Main Boundary Thrust regions. The XGBoost model successfully identified 78.06% of the validation landslides within these areas, achieving an efficiency ratio of 4.13. Vulnerability analysis suggests that around 87,000 residents (41% of the population) are at risk, along with 7,710 buildings (15.7%), 79 schools (24.2%), and 85 km of transport infrastructure. Validation against historical damage records showed about 78% agreement. These findings provide valuable spatial risk insights for planning and disaster mitigation in the data-scarce Himalayan regions, supporting evidence-based decision-making and policy implementation.