Efficient utilizing spatio-temporal information of consecutive frames is the core of lightweight video super-resolution. Many researchers employ alignment or setting hidden states to aggregate spatio-temporal information. However, these methods are somewhat crude in collecting and optimizing spatio-temporal features, which reduces the information utilization and reconfiguration capabilities of the network. Thus, to alleviate these problems, we propose a lightweight spatio-temporal attention network. We design a frame selection and spatio-temporal attention module, which can effectively filter, collect, and fuse inter-frame information. Moreover, the backward spatial fusion module and forward spatial fusion module are proposed to aggregate spatio-temporal information over long distances. Meanwhile, we design the spatial supplementation module to enhance the optimization and reconstruction capabilities of the network. Experiments indicate that our model achieves state-of-the-art performance on multiple datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight Spatio-Temporal Attention Network for Video Super-Resolution

  • Guofang Li,
  • Yonggui Zhu,
  • Zijun Zhao

摘要

Efficient utilizing spatio-temporal information of consecutive frames is the core of lightweight video super-resolution. Many researchers employ alignment or setting hidden states to aggregate spatio-temporal information. However, these methods are somewhat crude in collecting and optimizing spatio-temporal features, which reduces the information utilization and reconfiguration capabilities of the network. Thus, to alleviate these problems, we propose a lightweight spatio-temporal attention network. We design a frame selection and spatio-temporal attention module, which can effectively filter, collect, and fuse inter-frame information. Moreover, the backward spatial fusion module and forward spatial fusion module are proposed to aggregate spatio-temporal information over long distances. Meanwhile, we design the spatial supplementation module to enhance the optimization and reconstruction capabilities of the network. Experiments indicate that our model achieves state-of-the-art performance on multiple datasets.