<p>Speech emotion recognition (SER) is a critical component in domains such as human–computer interaction, mental health monitoring, and affective computing. While existing approaches—particularly those based on deep learning—have made significant progress, they often struggle to accurately capture subtle emotional cues in speech, especially in real-world scenarios. A key gap remains in the adaptability and decision-making capabilities of current models when faced with dynamic and ambiguous emotional expressions. To address this, we propose a novel speech emotion recognition framework that integrates deep reinforcement learning (DRL) with conventional classification models. Unlike static models, our approach enables adaptive learning and optimal decision policies for emotion classification over time. We evaluate our method using the widely adopted RAVDESS benchmark dataset. Experimental results demonstrate a notable improvement in performance, with our method achieving an accuracy of 89.5% and an <i>F</i>1 score of 0.86—outperforming several state-of-the-art baselines. These findings suggest that DRL offers a promising pathway toward more robust and context-aware speech emotion recognition systems.&#xa0;</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RELUEM-Reinforcing Emotional Understanding: Advancing Speech Emotion Recognition Through Deep Reinforcement Learning

  • Zhiping Zhang,
  • Rejuwan Shamim,
  • Anurag Sinha,
  • Pooja Jha,
  • Ahmed Alkhayyat

摘要

Speech emotion recognition (SER) is a critical component in domains such as human–computer interaction, mental health monitoring, and affective computing. While existing approaches—particularly those based on deep learning—have made significant progress, they often struggle to accurately capture subtle emotional cues in speech, especially in real-world scenarios. A key gap remains in the adaptability and decision-making capabilities of current models when faced with dynamic and ambiguous emotional expressions. To address this, we propose a novel speech emotion recognition framework that integrates deep reinforcement learning (DRL) with conventional classification models. Unlike static models, our approach enables adaptive learning and optimal decision policies for emotion classification over time. We evaluate our method using the widely adopted RAVDESS benchmark dataset. Experimental results demonstrate a notable improvement in performance, with our method achieving an accuracy of 89.5% and an F1 score of 0.86—outperforming several state-of-the-art baselines. These findings suggest that DRL offers a promising pathway toward more robust and context-aware speech emotion recognition systems.