Temporal Modeling via TCN and Transformer for Audio-Visual Emotion Recognition
摘要
Emotion recognition from multimodal sources is essential for advancing human-computer interaction. This paper introduces a comprehensive approach integrating robust methodologies to enhance multimodal emotion recognition. By harnessing both facial video and audio signals through Temporal Convolutional Networks (TCN) and Transformer models, alongside other advanced neural architectures, our ensemble effectively captures the complex dynamics of emotional expressions. Our approach has been rigorously evaluated on the Expression Classification challenge from the recent Affective Behavior Analysis in-the-Wild competition, demonstrating significant improvements compared to state-of-the-art methods. Specifically, our model achieved a 12% increase in F1 Score over the baseline, illustrating a substantial enhancement in performance and reliability.