Fine-Grained Spatial-Temporal Framework for Engagement Prediction
摘要
The engagement prediction task aims to identify the current level of involvement of individuals based on the information presented in video clips. In recent years, engagement prediction has attracted considerable attention due to its significance in real-life scenarios. We observed that an individual’s engagement status is often primarily manifested in the facial region. This observation prompted us to solve the engagement prediction problem by exploiting facial information to learn features related to engagement. In this paper, we propose a novel model for learning fine-grained features, called Fine-Grained Spatial-Temporal Framework (FGST). Specifically, we consider both spatial and temporal aspects. In the spatial domain, we propose an entropy-weighted method based on Fourier transform. By applying Fourier transform to video frames and introducing the concept of information entropy, we can segregate engagement-related information from engagement-unrelated information and mining difficult samples. This aids in enabling the model to learn features conducive to engagement prediction. In the temporal domain, we calculate the differences between adjacent frames and sum these differences to obtain regions of pixel variation between video frames, thereby generates saliency weight and improving model performance. Experiments on several public datasets demonstrate the proposed model can outperform the state-of-the-art methods.