<p>Multimodal Emotion Recognition in Conversation (MERC) has attracted significant attention in recent years, and existing methods mainly rely on contextual cues and multimodal interactions to predict emotions. However, these methods often suffer from the detrimental effects of noise, including contextual interaction noise and modality-specific noise contamination, leading to suboptimal model performance. Therefore, we propose DTDiff, a noise-aware framework that tackles two types of noise: inter-utterance interaction noise through an Adaptive Residual Decoupled Transformer (ARDT), and modality-specific noise via Language-Conditioned Denoising Learning (LDL). Specifically, ARDT improves robustness by effectively filtering irrelevant contextual dependencies and enhances the representation of each modality through decoupled residual attention fusion. Meanwhile, LDL employs language-conditioned diffusion models to denoise visual and acoustic modalities and measures noise levels via gating mechanisms. Finally, we introduce Dual-Signal Alignment to further promote multimodal fusion. Experiments on IEMOCAP and MELD datasets demonstrate that DTDiff outperforms state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DTDiff: adaptive decoupled transformer with language-conditioned denoising learning for multimodal emotion recognition in conversation

  • Tingting Zhang,
  • Xiaofei Zhu

摘要

Multimodal Emotion Recognition in Conversation (MERC) has attracted significant attention in recent years, and existing methods mainly rely on contextual cues and multimodal interactions to predict emotions. However, these methods often suffer from the detrimental effects of noise, including contextual interaction noise and modality-specific noise contamination, leading to suboptimal model performance. Therefore, we propose DTDiff, a noise-aware framework that tackles two types of noise: inter-utterance interaction noise through an Adaptive Residual Decoupled Transformer (ARDT), and modality-specific noise via Language-Conditioned Denoising Learning (LDL). Specifically, ARDT improves robustness by effectively filtering irrelevant contextual dependencies and enhances the representation of each modality through decoupled residual attention fusion. Meanwhile, LDL employs language-conditioned diffusion models to denoise visual and acoustic modalities and measures noise levels via gating mechanisms. Finally, we introduce Dual-Signal Alignment to further promote multimodal fusion. Experiments on IEMOCAP and MELD datasets demonstrate that DTDiff outperforms state-of-the-art methods.