The engagement prediction task aims to identify the current level of involvement of individuals based on the information presented in video clips. In recent years, engagement prediction has attracted considerable attention due to its significance in real-life scenarios. We observed that an individual’s engagement status is often primarily manifested in the facial region. This observation prompted us to solve the engagement prediction problem by exploiting facial information to learn features related to engagement. In this paper, we propose a novel model for learning fine-grained features, called Fine-Grained Spatial-Temporal Framework (FGST). Specifically, we consider both spatial and temporal aspects. In the spatial domain, we propose an entropy-weighted method based on Fourier transform. By applying Fourier transform to video frames and introducing the concept of information entropy, we can segregate engagement-related information from engagement-unrelated information and mining difficult samples. This aids in enabling the model to learn features conducive to engagement prediction. In the temporal domain, we calculate the differences between adjacent frames and sum these differences to obtain regions of pixel variation between video frames, thereby generates saliency weight and improving model performance. Experiments on several public datasets demonstrate the proposed model can outperform the state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Grained Spatial-Temporal Framework for Engagement Prediction

  • Shanshan Wang,
  • Yong Zheng,
  • Keyang Wang,
  • Xingyi Zhang

摘要

The engagement prediction task aims to identify the current level of involvement of individuals based on the information presented in video clips. In recent years, engagement prediction has attracted considerable attention due to its significance in real-life scenarios. We observed that an individual’s engagement status is often primarily manifested in the facial region. This observation prompted us to solve the engagement prediction problem by exploiting facial information to learn features related to engagement. In this paper, we propose a novel model for learning fine-grained features, called Fine-Grained Spatial-Temporal Framework (FGST). Specifically, we consider both spatial and temporal aspects. In the spatial domain, we propose an entropy-weighted method based on Fourier transform. By applying Fourier transform to video frames and introducing the concept of information entropy, we can segregate engagement-related information from engagement-unrelated information and mining difficult samples. This aids in enabling the model to learn features conducive to engagement prediction. In the temporal domain, we calculate the differences between adjacent frames and sum these differences to obtain regions of pixel variation between video frames, thereby generates saliency weight and improving model performance. Experiments on several public datasets demonstrate the proposed model can outperform the state-of-the-art methods.