Emotion recognition from multimodal sources is essential for advancing human-computer interaction. This paper introduces a comprehensive approach integrating robust methodologies to enhance multimodal emotion recognition. By harnessing both facial video and audio signals through Temporal Convolutional Networks (TCN) and Transformer models, alongside other advanced neural architectures, our ensemble effectively captures the complex dynamics of emotional expressions. Our approach has been rigorously evaluated on the Expression Classification challenge from the recent Affective Behavior Analysis in-the-Wild competition, demonstrating significant improvements compared to state-of-the-art methods. Specifically, our model achieved a 12% increase in F1 Score over the baseline, illustrating a substantial enhancement in performance and reliability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal Modeling via TCN and Transformer for Audio-Visual Emotion Recognition

  • Aleksei A. Andreev,
  • Andrey V. Savchenko

摘要

Emotion recognition from multimodal sources is essential for advancing human-computer interaction. This paper introduces a comprehensive approach integrating robust methodologies to enhance multimodal emotion recognition. By harnessing both facial video and audio signals through Temporal Convolutional Networks (TCN) and Transformer models, alongside other advanced neural architectures, our ensemble effectively captures the complex dynamics of emotional expressions. Our approach has been rigorously evaluated on the Expression Classification challenge from the recent Affective Behavior Analysis in-the-Wild competition, demonstrating significant improvements compared to state-of-the-art methods. Specifically, our model achieved a 12% increase in F1 Score over the baseline, illustrating a substantial enhancement in performance and reliability.