Multimodal entity linking (MEL) task aims to link ambiguous mentions in a multimodal context to entities in a multimodal knowledge base. Current methodologies predominantly employ intricate mechanisms to mitigate the discrepancies in modality. Nonetheless, these techniques often fall short of sufficiently distinguish entities with fine-grained semantic differences, particularly those classified under the same coarse-grained semantic class. In response, this paper introduces a Multimodal Entity Linking Framework with Large Language Models (LLMs) and Fine-Grained Semantic Class (FiGMEL). Specifically, we utilize LLMs in conjunction with triplet data from knowledge graphs to generate accurate and concise entity descriptions and to assign them fine-grained semantic classes. Then, after integrating modal features, FiGMEL discerns the nuances between entities within the same coarse-grained semantic class via three distinct pre-training exercises. Finally, based on the similarity between the mentions and entities, we generate the top-k candidate entities and use the LLMs for re-ranking. Experiments on three datasets show that our model outperforms existing models, and ablation studies affirm the effectiveness of our modules.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FiGMEL: A Multimodal Entity Linking Framework with Large Language Models and Fine-Grained Semantic Class

  • Chenkang Zhu,
  • Bohan Yu,
  • Pengcheng Wu,
  • Lei Zhuang,
  • Hongying Zan,
  • Kunli Zhang

摘要

Multimodal entity linking (MEL) task aims to link ambiguous mentions in a multimodal context to entities in a multimodal knowledge base. Current methodologies predominantly employ intricate mechanisms to mitigate the discrepancies in modality. Nonetheless, these techniques often fall short of sufficiently distinguish entities with fine-grained semantic differences, particularly those classified under the same coarse-grained semantic class. In response, this paper introduces a Multimodal Entity Linking Framework with Large Language Models (LLMs) and Fine-Grained Semantic Class (FiGMEL). Specifically, we utilize LLMs in conjunction with triplet data from knowledge graphs to generate accurate and concise entity descriptions and to assign them fine-grained semantic classes. Then, after integrating modal features, FiGMEL discerns the nuances between entities within the same coarse-grained semantic class via three distinct pre-training exercises. Finally, based on the similarity between the mentions and entities, we generate the top-k candidate entities and use the LLMs for re-ranking. Experiments on three datasets show that our model outperforms existing models, and ablation studies affirm the effectiveness of our modules.