Spatio-Temporal Scene Graph Reasoning Networks for Emotion Recognition in User-Generated Videos
摘要
This paper proposes a video emotion recognition framework based on visual relationship reasoning between objects and modal-aware fusion. To depict the emotional relationship between the region in several video frames between different regions in the same video frame, we create spatio-temporal scene graphs for object branch. Using the Graphormer to encode spatiotemporal scene graphs and quantify the emotional intensity between various objects, we introduce three structural information encoding methods. To achieve the goal of visual scene branch providing global context information for object branch, we propose a knowledge distillation mechanism based on object-aware. We use the Channel Temporal Attention Mechanism (CTAM) for audio streams to improve the characteristics of the information spectrum frames. Finally, we project the acoustic and visual features to two different subspaces to learn the features. The experimental results on Video Emotion-8 and Ekman-6 datasets prove the effectiveness of our proposed model.