A Prompt-Driven and Energy-Guided Framework for Multimodal Aspect-Based Sentiment Analysis
摘要
Multimodal Aspect-Based Sentiment Analysis (MABSA) has garnered significant attention in recent years due to its potential to identify aspect-specific sentiments within multimodal data. Despite the progress achieved by existing methods, several limitations remain: (i) the inability to effectively filter visual noise and the lack of task-specific feature alignment mechanisms; (ii) neglect of the semantic gap between text and image modalities, making it difficult to capture fine-grained cross-modal correlations; and (iii) existing span-based models fail to model the pairwise relationships between target span boundaries, resulting in unstable predictions. To address these issues, we propose an innovative framework, PEMSA, which incorporates two key modules: (1) Prompt-Driven Virtual Outlier Synthesis (PVOS), which dynamically generates virtual outliers to effectively filter sentiment foreground information and utilizes a prompt-guided mechanism to achieve cross-modal feature alignment; and (2) Energy-guided Contrastive Learning (EGCL), which employs energy-based modeling to capture the pairwise correlations between the start and end boundaries of target spans, thereby improving the stability and accuracy of boundary predictions. Experimental results on the benchmark datasets Twitter2015 and Twitter2017 demonstrate that PEMSA outperforms state-of-the-art methods across multiple evaluation metrics, showcasing its superior performance in multimodal sentiment analysis tasks.