<p>Multi-modal Entity Alignment (MMEA) aims to identify identical entities across different knowledge graphs that are linked to corresponding images. However, existing methods predominantly depend on pre-trained models for image modality processing, which often struggle to meet the fine visual information requirements essential for multi-modal alignment tasks. To address this issue, we propose a novel MMEA model, termed VEMEA, in which a new dual attention feature enhancement module was designed to enhance visual features extracted using pre-trained models. This module utilizes a multi-branch convolutional network alongside depth separable convolutional to enhance the local semantic information from the features output by the pre-trained model at first. Then the dual attention network dynamically adjusts the spatial and channel distribution of image features, capturing the global dependencies of spatially sensitive features and channels, thereby enhancing the representation of key regional features. Furthermore, during feature fusion, we employ a hierarchical semantic embedding and a cross-modal adaptive weighted attention mechanism to capture the correlations and interactions between modalities, thus alleviating issues related to modal heterogeneity. Additionally, this paper introduces an iterative optimization strategy based on credibility, which leverages the confidence scores of model outputs to preferentially select high-confidence aligned entity pairs, thereby increasing the number of marked entity pairs and improving alignment task performance. Finally, we conduct comprehensive experiments, and the results demonstrate that our model surpasses existing methods across all metrics, indicating its efficacy and robustness in the MMEA task.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-modal entity alignment method based on image enhancement and credibility iteration

  • Huayu Li,
  • Chi Zhang,
  • Xinxin Chen,
  • Cuicui Wang,
  • Yang Yue

摘要

Multi-modal Entity Alignment (MMEA) aims to identify identical entities across different knowledge graphs that are linked to corresponding images. However, existing methods predominantly depend on pre-trained models for image modality processing, which often struggle to meet the fine visual information requirements essential for multi-modal alignment tasks. To address this issue, we propose a novel MMEA model, termed VEMEA, in which a new dual attention feature enhancement module was designed to enhance visual features extracted using pre-trained models. This module utilizes a multi-branch convolutional network alongside depth separable convolutional to enhance the local semantic information from the features output by the pre-trained model at first. Then the dual attention network dynamically adjusts the spatial and channel distribution of image features, capturing the global dependencies of spatially sensitive features and channels, thereby enhancing the representation of key regional features. Furthermore, during feature fusion, we employ a hierarchical semantic embedding and a cross-modal adaptive weighted attention mechanism to capture the correlations and interactions between modalities, thus alleviating issues related to modal heterogeneity. Additionally, this paper introduces an iterative optimization strategy based on credibility, which leverages the confidence scores of model outputs to preferentially select high-confidence aligned entity pairs, thereby increasing the number of marked entity pairs and improving alignment task performance. Finally, we conduct comprehensive experiments, and the results demonstrate that our model surpasses existing methods across all metrics, indicating its efficacy and robustness in the MMEA task.