<p>This paper is about predictive learning, which is generating future frames given previous images. Suffering from the vanishing gradient problem, existing methods based on RNN and CNN can’t capture the long-term dependencies effectively. To overcome the above dilemma, we present MastNet a spatiotemporal framework for long-term predictive learning. In this paper, we design a Transformer-based encoder-decoder with hierarchical structure. As for the transformer block, we adopt the spatiotemporal window based self-attention to reduce computational complexity and the spatiotemporal shifted window partitioning approach. More importantly, we build a spatiotemporal autoencoder by the random clip mask strategy, which leads to better feature mining for temporal dependencies and spatial correlations. Furthermore, we insert an auxiliary prediction head, which can help our model generate higher-quality frames. Experimental results show that the proposed MastNet achieves the best results in accuracy and long-term prediction on two spatiotemporal datasets compared with the state-of-the-art models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A masked autoencoder network for spatiotemporal predictive learning

  • Fengzhen Sun,
  • Weidong Jin

摘要

This paper is about predictive learning, which is generating future frames given previous images. Suffering from the vanishing gradient problem, existing methods based on RNN and CNN can’t capture the long-term dependencies effectively. To overcome the above dilemma, we present MastNet a spatiotemporal framework for long-term predictive learning. In this paper, we design a Transformer-based encoder-decoder with hierarchical structure. As for the transformer block, we adopt the spatiotemporal window based self-attention to reduce computational complexity and the spatiotemporal shifted window partitioning approach. More importantly, we build a spatiotemporal autoencoder by the random clip mask strategy, which leads to better feature mining for temporal dependencies and spatial correlations. Furthermore, we insert an auxiliary prediction head, which can help our model generate higher-quality frames. Experimental results show that the proposed MastNet achieves the best results in accuracy and long-term prediction on two spatiotemporal datasets compared with the state-of-the-art models.