Speech Emotion Recognition (SER) holds a significant position in the fields of Natural Language Processing (NLP) and Affective Computing. Traditional unimodal approaches are constrained by the quality of speech signals and variations in speaker identity, thereby affecting the accuracy and comprehensiveness of emotion recognition. This paper proposes a novel multimodal speech emotion recognition model that fuses speech and text, named the Text and Speech Multimodal Emotion Fusion Model (TS-MEFM). The model introduces relative positional encoding and residual units through the Text Multimodal Attention SFKAN Encoder (TMAK), improving the encoder’s adaptability and robustness when processing utterances of varying lengths. By incorporating the Speech Multimodal Temporal SFKAN Encoder (SMTK) and the Speech Multimodal Attention Module (SMAM), the model more effectively captures subtle changes and transient features in speech signals while reducing computational overhead. Additionally, the Super Fast KAN (SFKAN) module enhances the model’s nonlinear modeling capacity and reduces parameter size. Experimental comparisons on the IEMOCAP and MELD datasets demonstrate that the proposed TS-MEFM model significantly outperforms the current state-of-the-art models in speech emotion recognition performance. Ablation studies further verify the contributions of each module to the overall performance of the model, proving the effectiveness of the proposed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TS-MEFM: A New Multimodal Speech Emotion Recognition Network Based on Speech and Text Fusion

  • Wei Wei,
  • Bingkun Zhang,
  • Yibing Wang

摘要

Speech Emotion Recognition (SER) holds a significant position in the fields of Natural Language Processing (NLP) and Affective Computing. Traditional unimodal approaches are constrained by the quality of speech signals and variations in speaker identity, thereby affecting the accuracy and comprehensiveness of emotion recognition. This paper proposes a novel multimodal speech emotion recognition model that fuses speech and text, named the Text and Speech Multimodal Emotion Fusion Model (TS-MEFM). The model introduces relative positional encoding and residual units through the Text Multimodal Attention SFKAN Encoder (TMAK), improving the encoder’s adaptability and robustness when processing utterances of varying lengths. By incorporating the Speech Multimodal Temporal SFKAN Encoder (SMTK) and the Speech Multimodal Attention Module (SMAM), the model more effectively captures subtle changes and transient features in speech signals while reducing computational overhead. Additionally, the Super Fast KAN (SFKAN) module enhances the model’s nonlinear modeling capacity and reduces parameter size. Experimental comparisons on the IEMOCAP and MELD datasets demonstrate that the proposed TS-MEFM model significantly outperforms the current state-of-the-art models in speech emotion recognition performance. Ablation studies further verify the contributions of each module to the overall performance of the model, proving the effectiveness of the proposed approach.