Transformer-Based Heterogeneous Feature Disentangled Representation Learning for Multimodal Sentiment Analysis
摘要
Multimodal Sentiment Analysis (MSA) aims to provide machines with multiple signals and train them to recognize the underlying emotions or intentions. However, previous studies have primarily focused on fusion strategies while overlooking the inherent heterogeneity between unimodal representations caused by significant distribution gaps. As a result, this disparity hinders the effective modeling of cross-modal emotional relations during fusion, ultimately degrading the model’s performance. To address this issue, we propose a shared transformer-based separation encoder to separate the raw intra-modal features into modality-specific tokens and modality-invariant tokens, and then reconstruct to obtain refined unimodal representations. We introduce multiple elaborated constraints to regulate this process. Unlike traditional disentanglement methods that rely on multiple encoders—resulting in high parameter costs and neglecting inter-modal interactions—our approach leverages a unified transformer encoder, which naturally facilitates cross-modal interaction learning while sharing parameters. Extensive experiments on widely-used multimodal datasets demonstrate that our approach outperforms existing methods and validates the effectiveness of our heterogeneous feature disentangling encoder.