Non-Local Spatial-Wise and Global Channel-Wise Transformer for Efficient Image Super-Resolution
摘要
Transformer-based methods have made favorable breakthroughs in image super-resolution (SR) due to the strong ability of capturing long-range dependencies in images. However, these methods mainly concentrate on capturing spatial interaction information, often ignoring to explore the global characteristics across the channel dimension. In this paper, we propose a novel Non-local Spatial-wise and Global Channel-wise Transformer (NSGCT) for efficient image SR. To comprehensively investigate inherent similarity information in both spatial and channel dimensions, we design a hybrid of Non-local Spatial-wise Self-Attention (NSSA) and Global Channel-wise Self-Attention (GCSA) within the Transformer layer. Specifically, NSSA is shifted-window-based and concentrates on the non-local spatial similarity features, while GCSA calculates the cross-covariance across the channels to exploit the global long-range image relationships. We also design an Efficient Gated Depth-wise-conv Feed-forward Network (EGDFN) as the feed-forward network to enhance and control the information flow in Transformer with an efficient implementation for further restoring the accurate texture information. Extensive quantitative and qualitative evaluations on benchmark datasets demonstrate that the proposed NSGCT performs favorably against other state-of-the-art efficient image SR methods in terms of computation costs and image reconstruction quality.