Enhancing Speech Emotion Recognition Using Transfer Learning from Speaker Embeddings
摘要
Understanding and identifying emotions from speech is a key challenge in automatic Speech Emotion Recognition (SER). Speech carries a variety of information about speaker’s emotional state or contextual emotions, but the lack of large and diverse emotional datasets makes it hard to apply advanced deep learning models for development of realiable and robust SER systems. Our study introduces a methodology that uses transfer learning and data augmentation to improve SER systems’ ability to classify emotional states accurately. Specifically, we focus on enhancing and assessing the performance of x-vector and r-vector speaker embedding models for SER task through pretraining the models on a large amount of speaker-labeled data following by fine-tuning on downstream emotional dataset. Testing the proposed approach on IEMOCAP and CREMA-D datasets shows notable increment in SER accuracy and thus usefulness of such cross-task transfer learning. Using transfer learning and data augmentation, our approach notably improved SER performance, achieving above 74% and 80% of accuracy on the IEMOCAP, and CREMA-D datasets, respectively.