Traditional video super-resolution (VSR) methods, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), often face challenges in maintaining temporal consistency and dealing with complex motion dynamics. This paper proposes an adaptive VSR method that combines an adaptive alignment mechanism with the Swin Transformer to improve video quality across various motion scenarios. By utilizing the self-attention capabilities of Transformers, the proposed method effectively captures both spatial and temporal dependencies. The key innovation of this approach lies in the adaptive alignment module, which dynamically selects the optimal alignment strategy—static, patch-based, or guided deformable attention (GDA) based on the level of motion between frames. This adaptability allows the model to flexibly handle different motion scenarios, significantly improving frame coherence and detail reconstruction. Experimental evaluations on the Vimeo-90 K dataset demonstrate that the proposed method outperforms traditional CNN/RNN frameworks and existing Transformer-based approaches in terms of Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Based Video Super-Resolution Algorithm with Adaptive Alignment Strategy Selection Methods

  • Lei Zhang,
  • Yujie Li,
  • Xiaoming Tao,
  • Nan Zhao,
  • Fang Cui,
  • Hengjiang Wang

摘要

Traditional video super-resolution (VSR) methods, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), often face challenges in maintaining temporal consistency and dealing with complex motion dynamics. This paper proposes an adaptive VSR method that combines an adaptive alignment mechanism with the Swin Transformer to improve video quality across various motion scenarios. By utilizing the self-attention capabilities of Transformers, the proposed method effectively captures both spatial and temporal dependencies. The key innovation of this approach lies in the adaptive alignment module, which dynamically selects the optimal alignment strategy—static, patch-based, or guided deformable attention (GDA) based on the level of motion between frames. This adaptability allows the model to flexibly handle different motion scenarios, significantly improving frame coherence and detail reconstruction. Experimental evaluations on the Vimeo-90 K dataset demonstrate that the proposed method outperforms traditional CNN/RNN frameworks and existing Transformer-based approaches in terms of Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM).