Utilizing Speaker Models and Topic Markers for Emotion Recognition in Dialogues
摘要
In modern speech emotion recognition (SER), improving upon state-of-the-art systems requires increasing sophistication of extracting discriminative information from training data. The key aspects of our research include utilizing automatic speaker verification (ASV) and topic detection models to enhance emotion recognition in multimodal (audio and text) dialogues. We conduct a comparative analysis of modern SER models and experiment with various models for text (RoBERTa, TodKat, COSMIC) and audio (WavLM, Wav2Vec2) modalities. We also employ attention-based fusion of the modalities and a BiGRU-based classification approach. The results show that our proposed model, which combines RoBERTa, TodKat, Wav2Vec2, WavLM, and attention-based fusion, achieves an F1-score of 0.657, outperforming the state-of-the-art systems using the same selection of modalities. This performance is only surpassed by models that utilize additional modalities beyond audio and text. We demonstrate the effectiveness of our approach in improving SER by leveraging advanced techniques for extracting discriminative information from multimodal training data. The incorporation of ASV and topic detection, along with the novel fusion and classification methods, contributes to the enhanced performance of our proposed model compared to the existing state-of-the-art SER systems.