<p>Landslides are among the primary geological hazards globally, presenting substantial risks to human lives, infrastructure, and natural ecosystems. A key challenge in landslide susceptibility assessments is the uncertainty in non-landslide sample selection, which can significantly affect the reliability of prediction results. This study addresses this issue by developing a landslide susceptibility prediction (LSP) framework for the Lin’an District of Hangzhou, Zhejiang Province, China, that integrates GeoDetector with ensemble learning algorithms, specifically random forest (RF) and extreme gradient boosting (XGBoost). Using GeoDetector, 12 dominant factors related to landslide occurrence, including elevation, slope, aspect, annual rainfall, and normalized difference vegetation index, are identified. To mitigate the uncertainty in non-landslide sample selection, <i>N</i> (<i>N</i> = 1, 10, 100, 500, 1000, 3000) non-landslide samples are randomly selected for model training. <i>N</i> sets of landslide susceptibility indices (LSIs) are calculated for each raster unit, and statistical analyses are conducted to quantify the uncertainty in LSIs associated with non-landslide sample selection. The results indicate that the <i>N</i> sets of LSIs for each raster unit exhibit an approximately normal distribution rather than remaining constant. This statistical behavior enables a more accurate quantification of the uncertainty associated with non-landslide sample selection. Furthermore, conducting multiple iterations of non-landslide sample selection markedly improves the stability and reliability of LSP predictions, whereas repeated sampling effectively mitigates the impact of rare misclassification events, as evidenced by confusion matrix evaluations. Specifically, the area under the receiver operating characteristic curve and overall accuracy for the RF/XGBoost models increase from 0.9137/0.9281 (<i>N</i> = 1) to 0.9399/0.9417 (<i>N</i> = 3000) and from 0.8592/0.8641 (<i>N</i> = 1) to 0.9369/0.9369 (<i>N</i> = 3000), respectively. Thus, this study provides a robust methodological framework for landslide susceptibility assessment, offering valuable insights for disaster risk management and mitigation strategies in landslide-prone regions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Landslide susceptibility prediction in Lin’an District, China, using ensemble learning with non-landslide sample uncertainty

  • Kai Chen,
  • Huolang Fang,
  • Jun Jiang

摘要

Landslides are among the primary geological hazards globally, presenting substantial risks to human lives, infrastructure, and natural ecosystems. A key challenge in landslide susceptibility assessments is the uncertainty in non-landslide sample selection, which can significantly affect the reliability of prediction results. This study addresses this issue by developing a landslide susceptibility prediction (LSP) framework for the Lin’an District of Hangzhou, Zhejiang Province, China, that integrates GeoDetector with ensemble learning algorithms, specifically random forest (RF) and extreme gradient boosting (XGBoost). Using GeoDetector, 12 dominant factors related to landslide occurrence, including elevation, slope, aspect, annual rainfall, and normalized difference vegetation index, are identified. To mitigate the uncertainty in non-landslide sample selection, N (N = 1, 10, 100, 500, 1000, 3000) non-landslide samples are randomly selected for model training. N sets of landslide susceptibility indices (LSIs) are calculated for each raster unit, and statistical analyses are conducted to quantify the uncertainty in LSIs associated with non-landslide sample selection. The results indicate that the N sets of LSIs for each raster unit exhibit an approximately normal distribution rather than remaining constant. This statistical behavior enables a more accurate quantification of the uncertainty associated with non-landslide sample selection. Furthermore, conducting multiple iterations of non-landslide sample selection markedly improves the stability and reliability of LSP predictions, whereas repeated sampling effectively mitigates the impact of rare misclassification events, as evidenced by confusion matrix evaluations. Specifically, the area under the receiver operating characteristic curve and overall accuracy for the RF/XGBoost models increase from 0.9137/0.9281 (N = 1) to 0.9399/0.9417 (N = 3000) and from 0.8592/0.8641 (N = 1) to 0.9369/0.9369 (N = 3000), respectively. Thus, this study provides a robust methodological framework for landslide susceptibility assessment, offering valuable insights for disaster risk management and mitigation strategies in landslide-prone regions.