Comparisons of machine learning models for landslide susceptibility mapping in the Jiuzhaigou earthquake-affected area, China
摘要
This study aimed to compare the prediction performance of single and ensemble machine learning (ML) models in terms of landslide susceptibility mapping in areas affected by the 2017 Jiuzhaigou earthquake. The single ML models selected were the logistic regression (LR) and naïve Bayes (NB) algorithms, and the selected ensemble ML models were the C4.5 decision tree (C4.5 DT), random forest (RF), light gradient boosting machine (LightGBM), extreme gradient boosting (XGBoost), RF coupled with information value (RF‒IV), and XGBoost‒IV. In total, 2482 landslides were identified and used to create training (75%) and validation (25%) datasets. Next, 11 landslide condition factors passed the multicollinearity evaluation. The feature importance between these conditioning factors was determined by SHAP analysis. The prediction capability of the eight models was validated and compared in terms of different statistical indices. The results revealed that the XGBoost model performed best (AUC = 0.910, ACC = 0.85, AP = 0.89, and k = 0.70), followed by XGBoost‒IV (AUC = 0.908), RF‒IV (AUC = 0.902), RF (AUC = 0.899), NB (AUC = 0.899), LightGBM (AUC = 0.897), LR (AUC = 0.875), and C4.5 DT (AUC = 0.838). These findings indicate that the statistically coupled model may negatively affect the accuracy and stability of single ML models. SHAP identified peak ground acceleration, lithology, and precipitation as key drivers of landslides, revealing a nonlinear relationship between feature variables and landslide prediction. This study provides a reference for predicting potential landslide hazard-prone areas and for explainable artificial intelligence research based on ML algorithms.