<p>Multimodal sentiment analysis is a longstanding and compelling research area that aims to accurately identify and classify human emotional states by integrating multiple perceptual modalities. However, the inherent differences across modalities lead to distribution disparities and overlapping information, posing significant challenges in effectively understanding and synthesizing multimodal content. To address these challenges, we propose a Decomposition-Diffusion Model (D<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(^2\)</EquationSource> </InlineEquation>M) to enhance multimodal sentiment analysis. Specifically, D<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(^2\)</EquationSource> </InlineEquation>M decomposes multimodal information into three components: modality-invariant representations (capturing common characteristics across modalities), modality-specific representations (capturing unique features of each modality), and corresponding residual noise components (enhancing model robustness). These components are then fused through a diffusion module, generating cross-modal content that not only integrates the original modalities’ features but also introduces new characteristics to facilitate inter-modal communication. Experimental results on benchmark datasets such as CMU-MOSI and CMU-MOSEI demonstrate that D<InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(^2\)</EquationSource> </InlineEquation>M achieves performance comparable to state-of-the-art methods, highlighting its ability to effectively capture and integrate information across different modalities, leading to more accurate sentiment predictions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

D\(^2\)M: Decomposition-diffusion model for enhanced multimodal sentiment analysis

  • Liansong Zong,
  • Wenzhe Xu,
  • Dongfeng Hu,
  • Jie Wang

摘要

Multimodal sentiment analysis is a longstanding and compelling research area that aims to accurately identify and classify human emotional states by integrating multiple perceptual modalities. However, the inherent differences across modalities lead to distribution disparities and overlapping information, posing significant challenges in effectively understanding and synthesizing multimodal content. To address these challenges, we propose a Decomposition-Diffusion Model (D \(^2\) M) to enhance multimodal sentiment analysis. Specifically, D \(^2\) M decomposes multimodal information into three components: modality-invariant representations (capturing common characteristics across modalities), modality-specific representations (capturing unique features of each modality), and corresponding residual noise components (enhancing model robustness). These components are then fused through a diffusion module, generating cross-modal content that not only integrates the original modalities’ features but also introduces new characteristics to facilitate inter-modal communication. Experimental results on benchmark datasets such as CMU-MOSI and CMU-MOSEI demonstrate that D \(^2\) M achieves performance comparable to state-of-the-art methods, highlighting its ability to effectively capture and integrate information across different modalities, leading to more accurate sentiment predictions.