Multimodal Named Entity Recognition (MNER) and Relation Extraction (MRE) are fundamental tasks in information extraction. However, existing approaches face challenges of the underutilization of diverse image representations, difficulties in aligning modalities due to inherent modality gaps, and the interference of modality noise. To address these challenges, we propose SAGE, a unified framework that combines Semantic Anchor and Granularity Enhancement. Specifically, SAGE effectively combines both pixel-level and semantic representations of images to complement the semantic deficiencies present in textual information, thereby enriching feature representation and improving contextual understanding. Then, a semantic anchor contrastive learning module is introduced to align image and text representations prior to modality fusion. To enhance cross-modal interactions and mitigate the impact of modality noise, we propose a multi-granularity visual-textual collaborative fusion module, which dynamically adjusts the visual information weights using text-guided gating mechanisms and models the relationships between candidate textual entities and image objects. Extensive experiments on three benchmark datasets demonstrate the superior performance of our SAGE, achieving state-of-the-art results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SAGE: A Unified Multimodal Entity and Relation Extraction Framework via Semantic Anchor and Granularity Enhancement

  • Jinkang Zheng,
  • Yahui Zhao,
  • Guozhe Jin,
  • Zhenguo Zhang,
  • Rongyi Cui

摘要

Multimodal Named Entity Recognition (MNER) and Relation Extraction (MRE) are fundamental tasks in information extraction. However, existing approaches face challenges of the underutilization of diverse image representations, difficulties in aligning modalities due to inherent modality gaps, and the interference of modality noise. To address these challenges, we propose SAGE, a unified framework that combines Semantic Anchor and Granularity Enhancement. Specifically, SAGE effectively combines both pixel-level and semantic representations of images to complement the semantic deficiencies present in textual information, thereby enriching feature representation and improving contextual understanding. Then, a semantic anchor contrastive learning module is introduced to align image and text representations prior to modality fusion. To enhance cross-modal interactions and mitigate the impact of modality noise, we propose a multi-granularity visual-textual collaborative fusion module, which dynamically adjusts the visual information weights using text-guided gating mechanisms and models the relationships between candidate textual entities and image objects. Extensive experiments on three benchmark datasets demonstrate the superior performance of our SAGE, achieving state-of-the-art results.