In video learning, accurately identifying learners’ emotional state is fundamental for enhancing the learner experience and improving teaching effectiveness. Most existing methods focus on analyzing learners’ physiological signals, neglecting the impact of instructional videos on learners, especially the hidden high-order relationships between video stimuli and physiological signals. This paper introduces a Multimodal Multi-Hpergraph emotion recognition framework, MMH, leveraging the strengths of hypergraphs in modeling high-order relationships. MMH captures high-order relationships between different modalities within instructional videos and between videos and learners, aiming to identify learners’ emotional states during video learning. During the hypergraph construction process, we employ two strategies to construct the hypergraph: (1) using multiple k-value feature KNN construction methods, and (2) construction methods based on vertex attributes. In the hypergraph learning process, we introduce a vertex attention mechanism to update vertex features during the hypergraph learning process, further improving the emotion recognition performance of MMH. Our method achieves state-of-the-art cross-subject emotion recognition performance on the VLMED dataset. Experimental results on two public benchmark datasets, MAHNOB-HCI and DEAP, further verify the effectiveness of our method in emotion recognition tasks. Our source code will be made publicly available upon acceptance of the paper.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modeling High-Order Relationships Between Human and Video for Emotion Recognition in Video Learning

  • Hanxu Ai,
  • Xiaomei Tao,
  • Xingbing Li,
  • Yanling Gan

摘要

In video learning, accurately identifying learners’ emotional state is fundamental for enhancing the learner experience and improving teaching effectiveness. Most existing methods focus on analyzing learners’ physiological signals, neglecting the impact of instructional videos on learners, especially the hidden high-order relationships between video stimuli and physiological signals. This paper introduces a Multimodal Multi-Hpergraph emotion recognition framework, MMH, leveraging the strengths of hypergraphs in modeling high-order relationships. MMH captures high-order relationships between different modalities within instructional videos and between videos and learners, aiming to identify learners’ emotional states during video learning. During the hypergraph construction process, we employ two strategies to construct the hypergraph: (1) using multiple k-value feature KNN construction methods, and (2) construction methods based on vertex attributes. In the hypergraph learning process, we introduce a vertex attention mechanism to update vertex features during the hypergraph learning process, further improving the emotion recognition performance of MMH. Our method achieves state-of-the-art cross-subject emotion recognition performance on the VLMED dataset. Experimental results on two public benchmark datasets, MAHNOB-HCI and DEAP, further verify the effectiveness of our method in emotion recognition tasks. Our source code will be made publicly available upon acceptance of the paper.