Two-stage landslide satellite image recognition in the southeastern tibet region based on Cascade R-CNN and SAM2
摘要
In this study, a two-stage landslide detection and segmentation framework based on Cascade R-CNN and Segment Anything Model 2 (SAM2) is proposed based on landslide features in Southeast Tibet. In the first stage, the Cascade R-CNN model trained with a transfer learning strategy achieves the initial detection of landslide areas, with a model average precision (mAP) of 82% on the Southeast Tibet landslide detection dataset. In the second phase, we constructed 9,000 landslide segmentation datasets containing augmented samples in Southeast Tibet, and achieved pixel-level segmentation by the bounding box adaptive extension technique and fine-tuned SAM2. The experimental results show that compared with benchmark models such as FCN (IoU = 84.4%), U-Net (IoU = 24.4%) and DeepLabv3 (IoU = 85.2%), the two-stage framework proposed in this paper achieves an intersection and concatenation ratio (IoU) of 94.3% and an overall accuracy of 96.1% on an independent test set, and the segmentation boundaries localization error (MAE) reduced by 0.26 percentage points compared to the original SAM2. Especially in challenging scenarios such as landslide areas with significant differences and vegetation covered areas, this method shows performance advantages over traditional algorithms. Through empirical analysis of representative landslide cases selected from Southeast Tibet, it is verified that the two-stage framework realizes sub-meter accuracy segmentation of landslide boundaries (1.07 m/pixel) by integrating the target detection and visual big model techniques, which provides reliable technical support for geohazard risk assessment in Southeast Tibet.