QSMT: Query-Shared Multimodal Transformer for Multimodal Sentiment Analysis
摘要
Transformer-based Multimodal Sentiment Analysis (MSA) has garnered growing attention for its strong ability to model cross-modal interaction and capture global features. Nevertheless, these methods generate substantial computational and GPU memory consumption, leading to model inefficiency and parameter redundancy. Therefore, we propose an efficient Query-Shared Multimodal Transformer (QSMT) for improving model efficiency and modeling robust cross-modal interaction. Specifically, we first propose a Query-Shared Cross-Attention (QSCA) mechanism, the core of QSMT, to fuse all modalities into multimodal fusion representation for interaction with unimodal representations, reducing computational complexity and space requirements of previous cross-modal interactions. Then, we further strengthen the supervisory role of the unimodal label generation module, which considers modality-specific information, on cross-modal interaction by applying average pooling. Extensive experimental results on commonly used benchmark datasets, including CMU-MOSI and CMU-MOSEI, demonstrate that the proposed QSMT achieves superior performance and significant efficiency gains.