<p>Recent research on enhancing image resolution using convolutional neural networks (CNNs) have shown encouraging outcomes. While due to the intrinsic locality of the convolution operator, CNN-based methods limit the capacity to obtain contextual information and long-range dependency. To address this problem, we propose a hybrid network by integrating CNN and Transformer which show impressive performance to learn long-range contextual information for image SR. Specifically, by introducing a spatial pyramid pooling (SPP) module into the multi-head attention (MHA), the Spatial Pyramid Swin Transformer (SPST) module achieves linear computational complexity and integrates multi-scale features. This enables the model to learn a wider range of multi-scale features and enhances the capabilities of the attention matrix. Moreover, the gated convolution (GC) module employs the abundant low-frequency information from low-resolution to assist reconstruction and provides a learnable dynamic feature selection mechanism to further constrain the training to improve performance. Extensive experiments were carried out to assess the efficacy of our approach utilizing a benchmark dataset. The results of indicate that our method surpasses alternative approaches in terms of parameter count and computational efficiency. Especially, the proposed method increases PSNR by 0.05 dB and uses 1.6M fewer parameters than SwinIR, resulting in a shorter inference time.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spstnet: image super-resolution using spatial pyramid swin transformer network

  • Yemei Sun,
  • Jiao Wang,
  • Yue Yang,
  • Yan Zhang

摘要

Recent research on enhancing image resolution using convolutional neural networks (CNNs) have shown encouraging outcomes. While due to the intrinsic locality of the convolution operator, CNN-based methods limit the capacity to obtain contextual information and long-range dependency. To address this problem, we propose a hybrid network by integrating CNN and Transformer which show impressive performance to learn long-range contextual information for image SR. Specifically, by introducing a spatial pyramid pooling (SPP) module into the multi-head attention (MHA), the Spatial Pyramid Swin Transformer (SPST) module achieves linear computational complexity and integrates multi-scale features. This enables the model to learn a wider range of multi-scale features and enhances the capabilities of the attention matrix. Moreover, the gated convolution (GC) module employs the abundant low-frequency information from low-resolution to assist reconstruction and provides a learnable dynamic feature selection mechanism to further constrain the training to improve performance. Extensive experiments were carried out to assess the efficacy of our approach utilizing a benchmark dataset. The results of indicate that our method surpasses alternative approaches in terms of parameter count and computational efficiency. Especially, the proposed method increases PSNR by 0.05 dB and uses 1.6M fewer parameters than SwinIR, resulting in a shorter inference time.