TSMDM-Net: A Speech Emotion Recognition Model Based on Multi-scale Time Series Dynamic Modeling
摘要
Speech Emotion Recognition (SER), as a core technology of emotional intelligence, aims to analyze emotional perception in human-computer interaction through acoustic features. In response to the limitations of traditional Transformer models in global attention noise sensitivity and insufficient multi-scale temporal modeling, this paper proposes a Multi-Scale Time Series Dynamic Modeling Network (TSMDM-Net). The network innovatively integrates dynamic dilated causal convolution and deformable attention mechanisms through the collaborative architecture of the Temporal Convolutional Encoder (TCE) and Local-Global Interaction Module (LGIM). TCE extracts multi-resolution temporal features using exponentially dilated convolution kernels, while LGIM adaptively adjusts local attention windows through a dynamic routing strategy and achieves complementary modeling of local acoustic events and global prosodic patterns through a cross-layer weight sharing mechanism. Experiments on the IEMOCAP and MELD datasets demonstrate that TSMDM-Net effectively captures multi-scale emotional features and significantly outperforms mainstream Transformer models and other advanced methods in terms of overall performance. Ablation experiments further validate the core contribution of the TCE and LGIM modules to the performance improvement.