Multimodal Aspect-Based Sentiment Analysis (MABSA) has garnered significant attention in recent years due to its potential to identify aspect-specific sentiments within multimodal data. Despite the progress achieved by existing methods, several limitations remain: (i) the inability to effectively filter visual noise and the lack of task-specific feature alignment mechanisms; (ii) neglect of the semantic gap between text and image modalities, making it difficult to capture fine-grained cross-modal correlations; and (iii) existing span-based models fail to model the pairwise relationships between target span boundaries, resulting in unstable predictions. To address these issues, we propose an innovative framework, PEMSA, which incorporates two key modules: (1) Prompt-Driven Virtual Outlier Synthesis (PVOS), which dynamically generates virtual outliers to effectively filter sentiment foreground information and utilizes a prompt-guided mechanism to achieve cross-modal feature alignment; and (2) Energy-guided Contrastive Learning (EGCL), which employs energy-based modeling to capture the pairwise correlations between the start and end boundaries of target spans, thereby improving the stability and accuracy of boundary predictions. Experimental results on the benchmark datasets Twitter2015 and Twitter2017 demonstrate that PEMSA outperforms state-of-the-art methods across multiple evaluation metrics, showcasing its superior performance in multimodal sentiment analysis tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Prompt-Driven and Energy-Guided Framework for Multimodal Aspect-Based Sentiment Analysis

  • Zuotao Fu,
  • Xuejiao Wan,
  • Guanghui He

摘要

Multimodal Aspect-Based Sentiment Analysis (MABSA) has garnered significant attention in recent years due to its potential to identify aspect-specific sentiments within multimodal data. Despite the progress achieved by existing methods, several limitations remain: (i) the inability to effectively filter visual noise and the lack of task-specific feature alignment mechanisms; (ii) neglect of the semantic gap between text and image modalities, making it difficult to capture fine-grained cross-modal correlations; and (iii) existing span-based models fail to model the pairwise relationships between target span boundaries, resulting in unstable predictions. To address these issues, we propose an innovative framework, PEMSA, which incorporates two key modules: (1) Prompt-Driven Virtual Outlier Synthesis (PVOS), which dynamically generates virtual outliers to effectively filter sentiment foreground information and utilizes a prompt-guided mechanism to achieve cross-modal feature alignment; and (2) Energy-guided Contrastive Learning (EGCL), which employs energy-based modeling to capture the pairwise correlations between the start and end boundaries of target spans, thereby improving the stability and accuracy of boundary predictions. Experimental results on the benchmark datasets Twitter2015 and Twitter2017 demonstrate that PEMSA outperforms state-of-the-art methods across multiple evaluation metrics, showcasing its superior performance in multimodal sentiment analysis tasks.