<p>Effectively managing temporal dependencies and capturing accurate spatial features are dual challenges in video prediction. Autoregressive models typically use shallow recursive networks with limited time-step horizons, while non-autoregressive models may overlook inherent temporal dependencies. To tackle these challenges, we introduce UST-SU, a U-shaped spatiotemporal simple unit network. UST-SU utilizes a multilayer interactive sampling architecture to extract challenging spatial features for shallow autoregressive networks. The fusion of multi-dimensional spatial features during image reconstruction preserves crucial region-specific information. Our spatiotemporal simple unit (ST-SU), focusing on encoding global spatiotemporal information, achieves a balance in subsequent memory flows. By discarding redundant information and incorporating early-stage details, ST-SU efficiently handles spatiotemporal sequence tasks, making it a suitable core unit for integration. Additionally, we propose a partial autoregressive strategy to broaden the temporal receptive field, preserving temporal dependencies discarded by non-autoregressive models. Across diverse video prediction benchmarks, UST-SU consistently outperforms previous state-of-the-art models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

UST-SU: a U-shaped video prediction network based on partial autoregression

  • Zhaojun Cui,
  • Wei Tian,
  • Fan Luo,
  • Qi Liu,
  • Shengqin Jiang

摘要

Effectively managing temporal dependencies and capturing accurate spatial features are dual challenges in video prediction. Autoregressive models typically use shallow recursive networks with limited time-step horizons, while non-autoregressive models may overlook inherent temporal dependencies. To tackle these challenges, we introduce UST-SU, a U-shaped spatiotemporal simple unit network. UST-SU utilizes a multilayer interactive sampling architecture to extract challenging spatial features for shallow autoregressive networks. The fusion of multi-dimensional spatial features during image reconstruction preserves crucial region-specific information. Our spatiotemporal simple unit (ST-SU), focusing on encoding global spatiotemporal information, achieves a balance in subsequent memory flows. By discarding redundant information and incorporating early-stage details, ST-SU efficiently handles spatiotemporal sequence tasks, making it a suitable core unit for integration. Additionally, we propose a partial autoregressive strategy to broaden the temporal receptive field, preserving temporal dependencies discarded by non-autoregressive models. Across diverse video prediction benchmarks, UST-SU consistently outperforms previous state-of-the-art models.