<p>Multimodal sentiment analysis (MSA) seeks to enhance sentiment recognition by synthesizing complementary data from textual, audio, and visual modalities. Nevertheless, current methodologies exhibit two significant shortcomings: uniform modal compression that neglects the disparate sentiment contributions of various modalities, resulting in either excessive or inadequate compression, and inadequate cross-modal interaction throughout the compression process. To tackle these difficulties, we offer the Gate-Controlled Information Bottleneck Cross-Modal Attention Network (GICA), a two-stage hierarchical framework that jointly optimizes adaptive modal compression and cross-modal information interaction. GICA implements a dual-gate bottleneck mechanism that adaptively modifies compression intensity according to the sentiment contribution of each modality, hence addressing the issues of over-compression and under-compression present in current information bottleneck techniques. Additionally, a cross-modal attention mechanism is incorporated into the compression pipeline to capture inter-modal relationships among text, audio, and visual information. To alleviate information loss resulting from compression, audio and visual elements are rebuilt and integrated with the compressed representations using an adaptive multimodal fusion layer. Comprehensive studies on the CMU-MOSI and CMU-MOSEI benchmarks reveal that GICA routinely surpasses state-of-the-art methodologies, attaining enhancements of 0.8% in correlation on CMU-MOSI and 0.7% in accuracy on CMU-MOSEI compared to the most robust baseline.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GICA: gate-controlled information bottleneck cross-modal attention network

  • Yuhao Qiang,
  • Junfeng Shen

摘要

Multimodal sentiment analysis (MSA) seeks to enhance sentiment recognition by synthesizing complementary data from textual, audio, and visual modalities. Nevertheless, current methodologies exhibit two significant shortcomings: uniform modal compression that neglects the disparate sentiment contributions of various modalities, resulting in either excessive or inadequate compression, and inadequate cross-modal interaction throughout the compression process. To tackle these difficulties, we offer the Gate-Controlled Information Bottleneck Cross-Modal Attention Network (GICA), a two-stage hierarchical framework that jointly optimizes adaptive modal compression and cross-modal information interaction. GICA implements a dual-gate bottleneck mechanism that adaptively modifies compression intensity according to the sentiment contribution of each modality, hence addressing the issues of over-compression and under-compression present in current information bottleneck techniques. Additionally, a cross-modal attention mechanism is incorporated into the compression pipeline to capture inter-modal relationships among text, audio, and visual information. To alleviate information loss resulting from compression, audio and visual elements are rebuilt and integrated with the compressed representations using an adaptive multimodal fusion layer. Comprehensive studies on the CMU-MOSI and CMU-MOSEI benchmarks reveal that GICA routinely surpasses state-of-the-art methodologies, attaining enhancements of 0.8% in correlation on CMU-MOSI and 0.7% in accuracy on CMU-MOSEI compared to the most robust baseline.