Multimodal Emotion Recognition from Audio and Visual Data
摘要
Multimodal emotion recognition, a crucial area of research and technology, integrates data from various sensory modalities to enhance emotion assessment accuracy. It addresses the complexity of automatically detecting emotions, which manifest through diverse channels of expression. The technology finds applications in multimedia retrieval and human–computer interactions. Deep neural networks, particularly Convolutional Neural Networks (CNNs), have shown promising results in this field. In this study, an emotion recognition system leveraging both auditory and visual modalities was developed, achieving an accuracy of 52% on the Multimodal Emotion Lines Dataset (MELD) dataset. The system comprises two CNN branches for processing audio and video signals, along with data preprocessing, augmentation, and feature extraction operations for both modalities. Effective management of atypical data points and capturing contextual information are emphasized as crucial aspects of this approach.