错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient and Lightweight Structure for Spatial-Temporal Feature Extraction in Video Super Resolution

  • Xiaonan He,
  • Yukun Xia,
  • Yuansong Qiao,
  • Brian Lee,
  • Yuhang Ye

摘要

Video Super Resolution (VSR) model based on deep convolutional neural network (CNN) uses multiple Low-Resolution (LR) frames as input and has a strong ability to recover High-Resolution (HR) frames and maintain video temporal information. However, to realize the above advantages, VSR must consider both spatial and temporal information to improve the perceived quality of the output video, leading to expensive operations such as cross-frame convolution. Therefore, how to balance the output video quality and computational cost is a worthy issue to be studied. To address the above problem, we propose an efficient and lightweight multi-scale 3D video super-resolution scheme that arranges 3D convolution features extraction blocks using a U-Net structure to achieve multi-scale feature extraction in both spatial and temporal dimensions. Quantitative and qualitative evaluation results on public video datasets show that compared to other simple cascaded spatial-temporal feature extraction structures, an U-Net structure achieves comparable texture details and temporal consistency while with a significant reduction in computation costs and latency.