TACST: Time-Aware Transformer for Robust Speech Emotion Recognition
摘要
Speech Emotion Recognition (SER) is an important research direction in the fields of human-computer interaction and affective computing. Effectively extracting emotional features from complex speech signals has always been a challenging task. This paper proposes a Time-Aware Convolutional Speech Transformer (TACST), which combines a Hybrid Convolutional Extractor (HCE) and a Time-Attention Module (TAM) to effectively improve emotion recognition performance. The HCE, by integrating Convolutional Neural Networks (CNN) and Multi-Head Attention, is able to capture both local and global features. The TAM introduces an adaptive attention mechanism along the time dimension, further enhancing the model’s ability to capture the global temporal dynamics of emotional features. Experimental results show that this method achieves significant improvements in various metrics, including Weighted Accuracy, Unweighted Accuracy, and F1 scores on datasets such as IEMOCAP, DAIC-WOZ, and MELD, outperforming existing SOTA models.