Visual audio and textual triplet fusion network for multi-modal sentiment analysis
摘要
In this work, we introduce a novel framework for multimodal sentiment analysis by presenting the cross-modal triplet fusion network (CTFN), a unified architecture designed to tackle the intrinsic complexities of multimodal data integration. Unlike traditional approaches that often handle modalities in isolation or rely on simple concatenation strategies, the CTFN architecture innovatively combines cross-modal interactions with a sophisticated fusion mechanism. At its core, the network leverages a triplet attention mechanism that dynamically aligns and refines emotional cues across textual, auditory, and visual modalities. This is further enhanced by the selective kernel module, which selectively emphasizes salient features within and across modalities through a self-attention mechanism, enabling a more nuanced and effective representation of multimodal sentiment cues. TThe CTFN framework was extensively validated on the CMU-MOSI and CMU-MOSEI datasets, demonstrating significant improvements in sentiment detection accuracy, surpassing existing models.