An Improved Self-Attention for Video Prediction
摘要
In recent years, there has been a growing focus on video prediction research using deep learning. This research field is expected to be applicable to a wide variety of domains, such as transportation engineering and meteorology. Currently, collision avoidance support based on real-time situational awareness using advanced driver-assistance systems is considered a promising tool for reducing traffic accidents. Although the performances of these systems improve daily, they have limitations. By leveraging future videos predicted through video prediction, collision avoidance support can be achieved more rapidly than in traditional systems. However, predicting videos with complex situational changes, such as videos from the dashboard cameras, remains challenging, resulting in blurred future videos. In particular, blurring is pronounced in regions with large temporal motions and is proportional to the time scale. In this paper, we propose an improved self-attention mechanism that enables the extraction of regions with large temporal motions. Moreover, we effectively introduce both improved and standard self-attention mechanisms into an RNN-based model. We targeted videos from a dashboard camera and aimed to reduce blurring in regions with large temporal motions to achieve high-quality video prediction. The proposed method outperforms other recent methods on video prediction from the dashboard camera videos and achieves high-quality video prediction by reducing blurring in regions with large temporal motions. Furthermore, ablation studies are conducted to demonstrate the effectiveness of utilizing improved and standard self-attention mechanisms.