Speech Emotion Recognition (SER), as a core technology of emotional intelligence, aims to analyze emotional perception in human-computer interaction through acoustic features. In response to the limitations of traditional Transformer models in global attention noise sensitivity and insufficient multi-scale temporal modeling, this paper proposes a Multi-Scale Time Series Dynamic Modeling Network (TSMDM-Net). The network innovatively integrates dynamic dilated causal convolution and deformable attention mechanisms through the collaborative architecture of the Temporal Convolutional Encoder (TCE) and Local-Global Interaction Module (LGIM). TCE extracts multi-resolution temporal features using exponentially dilated convolution kernels, while LGIM adaptively adjusts local attention windows through a dynamic routing strategy and achieves complementary modeling of local acoustic events and global prosodic patterns through a cross-layer weight sharing mechanism. Experiments on the IEMOCAP and MELD datasets demonstrate that TSMDM-Net effectively captures multi-scale emotional features and significantly outperforms mainstream Transformer models and other advanced methods in terms of overall performance. Ablation experiments further validate the core contribution of the TCE and LGIM modules to the performance improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TSMDM-Net: A Speech Emotion Recognition Model Based on Multi-scale Time Series Dynamic Modeling

  • Wei Wei,
  • Yibing Wang,
  • Bingkun Zhang,
  • Xiaodong Duan

摘要

Speech Emotion Recognition (SER), as a core technology of emotional intelligence, aims to analyze emotional perception in human-computer interaction through acoustic features. In response to the limitations of traditional Transformer models in global attention noise sensitivity and insufficient multi-scale temporal modeling, this paper proposes a Multi-Scale Time Series Dynamic Modeling Network (TSMDM-Net). The network innovatively integrates dynamic dilated causal convolution and deformable attention mechanisms through the collaborative architecture of the Temporal Convolutional Encoder (TCE) and Local-Global Interaction Module (LGIM). TCE extracts multi-resolution temporal features using exponentially dilated convolution kernels, while LGIM adaptively adjusts local attention windows through a dynamic routing strategy and achieves complementary modeling of local acoustic events and global prosodic patterns through a cross-layer weight sharing mechanism. Experiments on the IEMOCAP and MELD datasets demonstrate that TSMDM-Net effectively captures multi-scale emotional features and significantly outperforms mainstream Transformer models and other advanced methods in terms of overall performance. Ablation experiments further validate the core contribution of the TCE and LGIM modules to the performance improvement.