A Phoneme-Aware Multi-task Learning Framework with Dynamic Prioritization for Speech Emotion Recognition
摘要
In human-computer interaction (HCI), speech emotion recognition (SER) is a key technology that enables machines to detect emotions in voices for more effective communication. Recently, pre-trained speech representations have demonstrated significant potential for SER. However, due to the inherent variability in individual emotional expressions, models often focus more on identity information than emotional information within these representations. This can hinder the ability to fully leverage the capabilities of pre-trained representations for emotion recognition. To address this challenge, we propose a phoneme-aware multi-task learning framework tailored for SER. Specifically, the framework incorporates phoneme recognition as an auxiliary task and leverages the speaker-invariance of phoneme articulation to reduce the influence of identity information. Additionally, a dynamic task prioritization strategy has been customized to tackle the imbalance during the multi-task training process, leveraging the synergistic effects between tasks. To further enhance the utilization of pre-trained representations, we introduce a Squeeze-and-Excitation module to extract emotion-related information. Experimental results on the IEMOCAP and MELD datasets demonstrate that our approach achieves state-of-the-art performance.