In recent years, there has been a growing focus on video prediction research using deep learning. This research field is expected to be applicable to a wide variety of domains, such as transportation engineering and meteorology. Currently, collision avoidance support based on real-time situational awareness using advanced driver-assistance systems is considered a promising tool for reducing traffic accidents. Although the performances of these systems improve daily, they have limitations. By leveraging future videos predicted through video prediction, collision avoidance support can be achieved more rapidly than in traditional systems. However, predicting videos with complex situational changes, such as videos from the dashboard cameras, remains challenging, resulting in blurred future videos. In particular, blurring is pronounced in regions with large temporal motions and is proportional to the time scale. In this paper, we propose an improved self-attention mechanism that enables the extraction of regions with large temporal motions. Moreover, we effectively introduce both improved and standard self-attention mechanisms into an RNN-based model. We targeted videos from a dashboard camera and aimed to reduce blurring in regions with large temporal motions to achieve high-quality video prediction. The proposed method outperforms other recent methods on video prediction from the dashboard camera videos and achieves high-quality video prediction by reducing blurring in regions with large temporal motions. Furthermore, ablation studies are conducted to demonstrate the effectiveness of utilizing improved and standard self-attention mechanisms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Improved Self-Attention for Video Prediction

  • Ryosuke Hata,
  • Yoshihisa Shinozawa

摘要

In recent years, there has been a growing focus on video prediction research using deep learning. This research field is expected to be applicable to a wide variety of domains, such as transportation engineering and meteorology. Currently, collision avoidance support based on real-time situational awareness using advanced driver-assistance systems is considered a promising tool for reducing traffic accidents. Although the performances of these systems improve daily, they have limitations. By leveraging future videos predicted through video prediction, collision avoidance support can be achieved more rapidly than in traditional systems. However, predicting videos with complex situational changes, such as videos from the dashboard cameras, remains challenging, resulting in blurred future videos. In particular, blurring is pronounced in regions with large temporal motions and is proportional to the time scale. In this paper, we propose an improved self-attention mechanism that enables the extraction of regions with large temporal motions. Moreover, we effectively introduce both improved and standard self-attention mechanisms into an RNN-based model. We targeted videos from a dashboard camera and aimed to reduce blurring in regions with large temporal motions to achieve high-quality video prediction. The proposed method outperforms other recent methods on video prediction from the dashboard camera videos and achieves high-quality video prediction by reducing blurring in regions with large temporal motions. Furthermore, ablation studies are conducted to demonstrate the effectiveness of utilizing improved and standard self-attention mechanisms.