<p>Multimodal aspect-based sentiment analysis aims to extract aspects and identify their corresponding sentiment polarities from text-image pairs, which is a fine-grained sentiment analysis task. However, the traditional multimodal sentiment analysis method neglects the homogeneity and heterogeneity among the fine-grained modes. To address these limitations, this paper proposes a dual attention-based graph convolutional neural network for multimodal sentiment analysis (DAGCN). Specifically, from a local perspective, we employ two different types of attention mechanism to achieve fine-grained alignment between aspect terms and text-picture pair features, thus facilitating effective intramodal and multimodal information interaction. From a global perspective, we construct a comprehensive graph-structured relational matrix to explore deep-seated text-image correlations and utilize a graph convolutional network to perform feature fusion and node classification, ultimately deriving aspect-level sentiment polarities. Extensive experiments on two multimodal public datasets demonstrate that DAGCN achieves significant performance improvements over existing methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual attention-based graph convolutional neural network for multimodal sentiment analysis

  • Na Qu,
  • Long Yu,
  • Shengwei Tian,
  • Pusen Xia,
  • Chaoyue Wu

摘要

Multimodal aspect-based sentiment analysis aims to extract aspects and identify their corresponding sentiment polarities from text-image pairs, which is a fine-grained sentiment analysis task. However, the traditional multimodal sentiment analysis method neglects the homogeneity and heterogeneity among the fine-grained modes. To address these limitations, this paper proposes a dual attention-based graph convolutional neural network for multimodal sentiment analysis (DAGCN). Specifically, from a local perspective, we employ two different types of attention mechanism to achieve fine-grained alignment between aspect terms and text-picture pair features, thus facilitating effective intramodal and multimodal information interaction. From a global perspective, we construct a comprehensive graph-structured relational matrix to explore deep-seated text-image correlations and utilize a graph convolutional network to perform feature fusion and node classification, ultimately deriving aspect-level sentiment polarities. Extensive experiments on two multimodal public datasets demonstrate that DAGCN achieves significant performance improvements over existing methods.