Lightweight Spatio-Temporal Attention Network for Video Super-Resolution
摘要
Efficient utilizing spatio-temporal information of consecutive frames is the core of lightweight video super-resolution. Many researchers employ alignment or setting hidden states to aggregate spatio-temporal information. However, these methods are somewhat crude in collecting and optimizing spatio-temporal features, which reduces the information utilization and reconfiguration capabilities of the network. Thus, to alleviate these problems, we propose a lightweight spatio-temporal attention network. We design a frame selection and spatio-temporal attention module, which can effectively filter, collect, and fuse inter-frame information. Moreover, the backward spatial fusion module and forward spatial fusion module are proposed to aggregate spatio-temporal information over long distances. Meanwhile, we design the spatial supplementation module to enhance the optimization and reconstruction capabilities of the network. Experiments indicate that our model achieves state-of-the-art performance on multiple datasets.