There are complex emotional interactions between individuals in group and between group and individuals. Although existing methods for group emotion recognition (GER) made quite efforts to learn the spatial-temporal characteristics via efficient deep learning-based networks, they neglected to model the interactive characteristics within group videos. In this article, we propose a graph-based GER approach to learn the spatial-temporal interactive characteristics from the facial and holistic cues of a group video. Specifically, we construct a spatial-temporal graph with facial information to describe the emotional relationships within a group in spatial and temporal dimensions. We employ a graph attention network (GAT) to dynamically model the emotional relationships and influences between individual and group nodes across the spatial-temporal dimension. The proposed method utilizes the GAT to explore the temporal correlations of holistic features extracted from the video frames. The introduced graph attention mechanism helps the proposed network effectively focus on the important nodes, capture interactive information, and generate a more precise spatial-temporal representation for GER. We fuse the decisions based on facial and holistic information in a linear way to obtain a comprehensive recognition result for the emotional state of group videos. Extensive experiments demonstrate that the proposed method learns effective spatial-temporal emotional features, and achieves superior performance in overall accuracies of 70.23% and 92.90% on the VGAF and GECV datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Spatial-Temporal Graph Convolutional Network for Video-Based Group Emotion Recognition

  • Xingzhi Wang,
  • Tao Chen,
  • Dong Zhang

摘要

There are complex emotional interactions between individuals in group and between group and individuals. Although existing methods for group emotion recognition (GER) made quite efforts to learn the spatial-temporal characteristics via efficient deep learning-based networks, they neglected to model the interactive characteristics within group videos. In this article, we propose a graph-based GER approach to learn the spatial-temporal interactive characteristics from the facial and holistic cues of a group video. Specifically, we construct a spatial-temporal graph with facial information to describe the emotional relationships within a group in spatial and temporal dimensions. We employ a graph attention network (GAT) to dynamically model the emotional relationships and influences between individual and group nodes across the spatial-temporal dimension. The proposed method utilizes the GAT to explore the temporal correlations of holistic features extracted from the video frames. The introduced graph attention mechanism helps the proposed network effectively focus on the important nodes, capture interactive information, and generate a more precise spatial-temporal representation for GER. We fuse the decisions based on facial and holistic information in a linear way to obtain a comprehensive recognition result for the emotional state of group videos. Extensive experiments demonstrate that the proposed method learns effective spatial-temporal emotional features, and achieves superior performance in overall accuracies of 70.23% and 92.90% on the VGAF and GECV datasets, respectively.