A disentanglement mamba network with a temporally slack reconstruction mechanism for multimodal continuous emotion recognition
摘要
While multimodal signals provide complementary information, they also introduce redundant information, hindering the learning of valuable information. Existing works on emotion recognition have achieved notable progress in multimodal fusion. However, the presence of redundant information remains one of the major challenges in this area. In this paper, we propose a Disentanglement Mamba network with a Temporally Slack Reconstruction Mechanism (DMA-TSM) to mitigate the impact of redundant information. The DMA-TSM employs an encoder-decoder framework. During the encoding phase, it decouples the original multimodal features into modality-common and modality-specific features. With explicit informational properties, these features inherently reduce redundant information at the feature level and alleviate the distributional gap between modalities. In the decoding phase, we present the TSM to avoid redundant information at the temporal level. By imposing a relaxed contextual constraint during the decoding and reconstruction process, the TSM balances preserving contextual information and avoiding redundant information at the temporal level. Additionally, by integrating Mamba with attention mechanisms, the DMA-TSM achieves substantial performance improvements with a relatively low computational cost. Experimental results on the Remote Collaborative and Affective (RECOLA) and the Ulm-Trier Social Stress Test (ULM-TSST) datasets show that our method is superior to the recently proposed methods, emphasizing the importance of reducing redundancy.