Knowledge bases (KBs) like Wikidata, while extensive, often have significant gaps, particularly for long-tail entities that are underrepresented but critical for comprehensive knowledge graphs. Language models (LMs) have emerged as a potential solution, yet existing methods largely focus on well-covered, prominent entities, overlooking the challenges of long-tail entities and unseen relations. To address this, we propose a novel knowledge base completion (KBC) approach that integrates zero-shot learning to handle entities and relations not seen during training. Our two-stage framework with zero-shot learning leverages one LM for candidate retrieval and another one for candidate verification and disambiguation, ensuring accurate and relevant fact completion. We used MALT, a dataset derived from Wikidata, to evaluate our method on frequent and long-tail entities. Our approach achieves state-of-the-art performance, with significant improvements in recall, demonstrating its effectiveness in bridging critical knowledge gaps and adapting to the dynamic nature of evolving KBs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Knowledge Base Completion for Long-Tail Entities Through Zero-Shot Learning

  • Sravani Mall,
  • Sayed Reza,
  • Thien Khai Tran,
  • Hien Nguyen

摘要

Knowledge bases (KBs) like Wikidata, while extensive, often have significant gaps, particularly for long-tail entities that are underrepresented but critical for comprehensive knowledge graphs. Language models (LMs) have emerged as a potential solution, yet existing methods largely focus on well-covered, prominent entities, overlooking the challenges of long-tail entities and unseen relations. To address this, we propose a novel knowledge base completion (KBC) approach that integrates zero-shot learning to handle entities and relations not seen during training. Our two-stage framework with zero-shot learning leverages one LM for candidate retrieval and another one for candidate verification and disambiguation, ensuring accurate and relevant fact completion. We used MALT, a dataset derived from Wikidata, to evaluate our method on frequent and long-tail entities. Our approach achieves state-of-the-art performance, with significant improvements in recall, demonstrating its effectiveness in bridging critical knowledge gaps and adapting to the dynamic nature of evolving KBs.