<p>Speech Emotion Recognition (SER) plays a pivotal role in advancing human-computer interaction. Existing SER methods, however, use a single type of feature as input or perform simple fusion of multiple features, which not only results in insufficient feature representation but also leads to the loss of critical information. To maximize the potential of the features, a multi-branch interactive attention network based on self-distillation, abbreviated as MIA-SD, is proposed. Firstly, three parallel branches are proposed to handle different types of speech features: sLSTM is innovatively introduced to capture Mel-frequency cepstral coefficients features, AlexNet is employed to process spectrogram features, and Wav2Vec2 is used to process raw audio for embedded high-level acoustic information. This multi-branch approach effectively compensates for the shortcomings of using a single type of feature in expressing emotional information. Secondly, a multi-branch interactive attention (MIA) mechanism is proposed to integrate multiple speech features. MIA establishes cross-feature interaction through shared key and value and improves feature fusion quality by using an attention mechanism, thereby improving SER performance. Thirdly, to prevent MIA from potential information loss or poor fusion quality during direct training, a training approach of self-distillation from triple-teacher to single-student is adopted. Specifically, the three parallel branches are used as teacher models to supervise the final output of the model. On the datasets IEMOCAP, EMO-DB, and CASIA, under both speaker-dependent and speaker-independent experimental settings, the strong performances of MIA-SD verify its effectiveness and performance advantages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-branch Interactive Attention Network Based on Self-Distillation for Speech Emotion Recognition

  • Yuanyuan Wei,
  • Heming Huang,
  • Kedi Huang,
  • Yonghong Fan,
  • Jie Zhou

摘要

Speech Emotion Recognition (SER) plays a pivotal role in advancing human-computer interaction. Existing SER methods, however, use a single type of feature as input or perform simple fusion of multiple features, which not only results in insufficient feature representation but also leads to the loss of critical information. To maximize the potential of the features, a multi-branch interactive attention network based on self-distillation, abbreviated as MIA-SD, is proposed. Firstly, three parallel branches are proposed to handle different types of speech features: sLSTM is innovatively introduced to capture Mel-frequency cepstral coefficients features, AlexNet is employed to process spectrogram features, and Wav2Vec2 is used to process raw audio for embedded high-level acoustic information. This multi-branch approach effectively compensates for the shortcomings of using a single type of feature in expressing emotional information. Secondly, a multi-branch interactive attention (MIA) mechanism is proposed to integrate multiple speech features. MIA establishes cross-feature interaction through shared key and value and improves feature fusion quality by using an attention mechanism, thereby improving SER performance. Thirdly, to prevent MIA from potential information loss or poor fusion quality during direct training, a training approach of self-distillation from triple-teacher to single-student is adopted. Specifically, the three parallel branches are used as teacher models to supervise the final output of the model. On the datasets IEMOCAP, EMO-DB, and CASIA, under both speaker-dependent and speaker-independent experimental settings, the strong performances of MIA-SD verify its effectiveness and performance advantages.