错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

WESER: Wav2Vec 2.0 Enhanced Speech Emotion Recognizer

  • Ahmed Ba Matraf,
  • Ashraf Elnagar

摘要

Emotion recognition is essential in human interactions, yet it is challenging due to cultural, linguistic, and individual differences. Advancements in deep learning technologies have effectively addressed these complexities. This paper proposes the WESER model, utilizing the Wav2Vec 2.0 model for Speech Emotion Recognition (SER) on RAVDESS and EMODB datasets. The XLSR variant of Wav2Vec 2.0 is pre-trained on a large collection of linguistic data, making it ideal for tasks with limited datasets. In our methodology, we first standardized audio samples for consistency and then performed fine-tuning. The results were notable, with WESER achieving a speaker-independent (SI) accuracy of 89.24% on RAVDESS and 97.20% on EMODB. However, the speaker-dependent (SD) accuracy was lower, particularly on RAVDESS at 73.33%, highlighting the challenges of adapting to individual speech patterns. The model’s efficiency is validated against existing models, demonstrating significant improvements in emotion recognition from speech.