Change detection (CD) has been widely utilized in various domains, including natural resource monitoring and urban construction management. However, recent advancements in change detection have primarily concentrated on feature fusion strategies, often neglecting intricate details. Consequently, this oversight frequently results in the omission of small targets in change detection outcomes and the emergence of region boundary connections. Res-Former is a transformer-based change detection network that aims to address the challenges associated with change detection tasks. Res-Former capitalizes on convolutional neural networks (CNNs), employing ResNet50 as the backbone for extracting high-level semantic features. Notably, we exclusively utilize the output from the first stage to mitigate the loss of spatial details and alleviate the issue of missing small targets. Additionally, depth-wise convolutions are integrated into the transformer encoder to reduce parameter count. By self-attention mechanisms within the encoder, we achieve finer edge segmentation and effectively eliminate boundary connections. Lastly, a lightweight multi-layer perceptron (MLP) decoder is employed for pixel-level prediction. Experiments on two benchmark datasets verify the effectiveness of our proposed network.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Transformer-Based Network for Remote Sensing Image Change Detection

  • Xiaoyi Cai,
  • Yuyu Tian,
  • Pengdao Xu

摘要

Change detection (CD) has been widely utilized in various domains, including natural resource monitoring and urban construction management. However, recent advancements in change detection have primarily concentrated on feature fusion strategies, often neglecting intricate details. Consequently, this oversight frequently results in the omission of small targets in change detection outcomes and the emergence of region boundary connections. Res-Former is a transformer-based change detection network that aims to address the challenges associated with change detection tasks. Res-Former capitalizes on convolutional neural networks (CNNs), employing ResNet50 as the backbone for extracting high-level semantic features. Notably, we exclusively utilize the output from the first stage to mitigate the loss of spatial details and alleviate the issue of missing small targets. Additionally, depth-wise convolutions are integrated into the transformer encoder to reduce parameter count. By self-attention mechanisms within the encoder, we achieve finer edge segmentation and effectively eliminate boundary connections. Lastly, a lightweight multi-layer perceptron (MLP) decoder is employed for pixel-level prediction. Experiments on two benchmark datasets verify the effectiveness of our proposed network.