Speech emotion recognition (SER) aims to identify and comprehend human emotions conveyed through speech, and has garnered widespread attention in recent years. Existing SER methods usually focus on extracting semantic features to discern emotions. However, these approaches often neglect additional speech characteristics, such as emotional nuances, limiting the efficacy of SER systems in real-world applications. Indeed, these emotional features are vital for SER tasks, as they complement semantic features by capturing the underlying feelings conveyed in speech. To this end, this paper is dedicated to the integration of both semantic and emotional features in speech, aiming to mitigate the loss of critical information. We propose a novel LightGBM-stacking architecture that employs an explicit feature fusion strategy to amalgamate semantic and emotional features derived from pre-trained models. We evaluate our model using multiple datasets, including IEMOCAP, RAVDESS-ACTOR, and RAVDESS-SONG. Experimental results show that the proposed LightGBM-stacking architecture significantly outperforms individual base classifiers. Moreover, when compared with the current best-known model, our architecture exhibits comparable or superior effectiveness, underscoring the benefits of feature fusion in enhancing SER systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bridging Semantic and Emotional Gaps in Speech Emotion Recognition: A Novel LightGBM-Stacking Architecture

  • Zhilong Duan,
  • Lin Zhou,
  • Degen Huang,
  • Xuewen Shi

摘要

Speech emotion recognition (SER) aims to identify and comprehend human emotions conveyed through speech, and has garnered widespread attention in recent years. Existing SER methods usually focus on extracting semantic features to discern emotions. However, these approaches often neglect additional speech characteristics, such as emotional nuances, limiting the efficacy of SER systems in real-world applications. Indeed, these emotional features are vital for SER tasks, as they complement semantic features by capturing the underlying feelings conveyed in speech. To this end, this paper is dedicated to the integration of both semantic and emotional features in speech, aiming to mitigate the loss of critical information. We propose a novel LightGBM-stacking architecture that employs an explicit feature fusion strategy to amalgamate semantic and emotional features derived from pre-trained models. We evaluate our model using multiple datasets, including IEMOCAP, RAVDESS-ACTOR, and RAVDESS-SONG. Experimental results show that the proposed LightGBM-stacking architecture significantly outperforms individual base classifiers. Moreover, when compared with the current best-known model, our architecture exhibits comparable or superior effectiveness, underscoring the benefits of feature fusion in enhancing SER systems.