A Global-Local Interactive Multimodal Sentiment Analysis Framework with Adaptive Attention and Multi-scale Feature Enhancement
摘要
Multimodal sentiment analysis faces dual challenges of asynchronous audio-visual fusion and robustness to modality absence. To address these, this paper proposes a Hierarchical Transformer with Dual-level Feature Enhancement (HTMD). The framework first employs mutual promotion units to dynamically calibrate asynchronous audio-visual streams: audio features are optimized through an adaptive time-frequency attention mechanism, while video representations are enhanced via a multi-scale local-global attention mechanism. For modality absence issues, the framework implements feature reconstruction complemented by a modified MAE loss function to ensure training stability. Comprehensive evaluations on CMU-MOSI and CH-SIMS datasets demonstrate that HTMD significantly outperforms existing methods in both complete and partial modality scenarios, establishing new benchmarks for robust multimodal sentiment analysis.