In this study, we conduct a comprehensive comparative analysis of various deep learning techniques for the Speech Emotion Recognition (SER) task, ranging from LSTM to CNN-BiLSTM-Multihead-Attention. Our evaluation utilizes the CREMA-D dataset for English and a crowdsourced dataset for the Algerian dialect. We propose two novel transformer-based frameworks, including an unimodal Transformer model and a multimodal Transformer model integrating DziriBERT. The Transformer-based model achieved an accuracy of 95% on the CREMA-D dataset, while the multimodal Transformer model with DziriBERT attained an impressive 97.42% accuracy on the Algerian dialect dataset. The proposed frameworks, unimodal Transformer and multimodal Transformer+DziriBERT significantly outperformed previous studies that used the same employed datasets in English language and Algerian dialect, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Speech Emotion Recognition: A Comparative Study and Enhanced Performance in English and Algerian Dialect Using Uni and Multimodal Transformer

  • Sabrina Khemis,
  • Abdelouahab Moussaoui,
  • Skander Hamdi

摘要

In this study, we conduct a comprehensive comparative analysis of various deep learning techniques for the Speech Emotion Recognition (SER) task, ranging from LSTM to CNN-BiLSTM-Multihead-Attention. Our evaluation utilizes the CREMA-D dataset for English and a crowdsourced dataset for the Algerian dialect. We propose two novel transformer-based frameworks, including an unimodal Transformer model and a multimodal Transformer model integrating DziriBERT. The Transformer-based model achieved an accuracy of 95% on the CREMA-D dataset, while the multimodal Transformer model with DziriBERT attained an impressive 97.42% accuracy on the Algerian dialect dataset. The proposed frameworks, unimodal Transformer and multimodal Transformer+DziriBERT significantly outperformed previous studies that used the same employed datasets in English language and Algerian dialect, respectively.