Simple Attention Deep Markov Model: A Scanpath Prediction Method Based on Three-Dimensional Attention Weights for 360° Images
摘要
Scanpath prediction for 360-degree images aims to simulate the sequence of fixation points of human observers on staring images to apply it to saliency detection and image quality assessment. Previous research relies on static saliency maps so lacks strong visual feature representation capabilities. In this study, we have developed a 360-degree image scanpath prediction method based on three-dimensional attention weights. We design a Simple Attention module, namely SimAtt, to infer the three-dimensional attention weights of 360-degree image feature mapping to accurately capture different areas in the panorama. We used the Deep Markov Model (DMM) as the main framework. Firstly, the starter initializes the scanpath’s starting point; then, the semantic-guided transition function controls the dynamics of the visual state, with the SimAtt module calculating the three-dimensional attention weight. Images and coordinates are learned using a sphere convolutional neural network and CoordConv layer, based on the importance of different regions in the image. Next, the emission function generates gaze points from the parameterized visual state. Finally, variational inference verifies the gaze point distribution. The 360-degree image prediction method combined with the SimAtt module presents satisfactory results, with the scanpath similarity measurement Recurrence (REC) reaching 3.329.