错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RCSTNet: Integrating Convolutional and Transformer Features for Semantic Segmentation of Remote Sensing Images

  • Tianyi Wang,
  • Zhongyun Liu

摘要

Semantic segmentation in high-resolution remote sensing (HR-RS) enables precise land cover classification by labeling every pixel with a designated category. However, conventional image cropping techniques used in neural network training often restrict the perception of extensive spatial dependencies. To overcome this challenge, we introduce RSCTNet, a fusion framework that integrates ResNet and Vision Transformer (ViT) to enhance semantic segmentation in HR RSIs. The model employs a dual-branch design: (1) a CNN-based branch for fine-grained local feature extraction, and (2) a dedicated Context Transformer Module that models long-range dependencies through self-attention, enhancing the discrimination of complex land-cover and land-use (LCLU) patterns. The Vaihingen dataset validation shows that RSCTNet surpasses current methods in both segmentation accuracy and computational performance. The importance of all building blocks was verified experimentally, particularly highlighting global context modeling's crucial role. Our work not only advances HR-RSI segmentation but also provides a scalable framework for integrating CNN and Transformer architectures in remote sensing applications.