<p>An increasing number of deep learning-based single image super-resolution (SISR) methods enhance reconstruction performances by assembling more complex modules or deepening network architecture, which brings a huge number of network parameters. Concurrently, these approaches often lack the interpretation for their network design. To deal with this issue, a novel interpretable Transformer-based network (ITN) is proposed based on optimization theory in this paper. Specifically, the proposed ITN method derives an algorithm from the image degradation optimization, which provides a theoretical explanation for network design. Then based on the above algorithm, an effective deep network is designed via a hybrid of MSRB and Transformer as a denoiser in iterative algorithm. As the proposed ITN is drawn in line with the iteration solution of the optimization problem, it has an algorithm interpretation and realizes the lightweight structure because of its form of recursion. Meanwhile, a depthwise separable convolution is used to calculate the cross-channel covariance matrix for implementing the self-attention mechanism in Transformer. Extensive experiments demonstrate that the proposed ITN outperforms existing related super-resolution methods in both reconstruction performance and architectural efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Super-Resolution via Interpretable Lightweight Transformers

  • Jianwei Zhao,
  • Tingwei Wang,
  • Jieyu Liu,
  • Yujie Zhu,
  • Zhenghua Zhou

摘要

An increasing number of deep learning-based single image super-resolution (SISR) methods enhance reconstruction performances by assembling more complex modules or deepening network architecture, which brings a huge number of network parameters. Concurrently, these approaches often lack the interpretation for their network design. To deal with this issue, a novel interpretable Transformer-based network (ITN) is proposed based on optimization theory in this paper. Specifically, the proposed ITN method derives an algorithm from the image degradation optimization, which provides a theoretical explanation for network design. Then based on the above algorithm, an effective deep network is designed via a hybrid of MSRB and Transformer as a denoiser in iterative algorithm. As the proposed ITN is drawn in line with the iteration solution of the optimization problem, it has an algorithm interpretation and realizes the lightweight structure because of its form of recursion. Meanwhile, a depthwise separable convolution is used to calculate the cross-channel covariance matrix for implementing the self-attention mechanism in Transformer. Extensive experiments demonstrate that the proposed ITN outperforms existing related super-resolution methods in both reconstruction performance and architectural efficiency.