<p>Knowledge graph completion (KGC) aims to predict missing facts in knowledge graphs(KGs). However, KGC faces two significant challenges: (i) low-quality entity descriptions in KGs and (ii) the structural sparsity of KGs. To address these issues, we propose a textual and structural dual enhancement (TSDE) framework that leverages large language models (LLMs) for graph data augmentation, thereby improving the performance of KGC models. To overcome the problem of low-quality entity descriptions, we designed a sampling algorithm to capture multiple paths surrounding entities, transforming this path information into text and integrating it into our carefully designed prompt templates to guide LLMs in generating high-quality descriptions enriched with contextual semantics. Additionally, to address the structural sparsity of KGs, we represent each entity description as a vector and calculate cosine similarity to identify a set of highly similar entities for each entity. We then assign similarity relations to these highly similar entity pairs to generate additional similar triples. After sampling these newly generated similar triples at a specified ratio, we incorporate them into the training set. We validated the effectiveness of the TSDE framework on two publicly available datasets: FB15k-237 and WN18RR. Experimental results demonstrate that the TSDE framework significantly enhances KGC model performance by refining entity descriptions and enriching KG structures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Textual and structural dual enhancement for knowledge graph completion with large language models

  • Liqin Wang,
  • Yifan Gan,
  • Xu Wang,
  • Yongfeng Dong,
  • Zhihong Xu

摘要

Knowledge graph completion (KGC) aims to predict missing facts in knowledge graphs(KGs). However, KGC faces two significant challenges: (i) low-quality entity descriptions in KGs and (ii) the structural sparsity of KGs. To address these issues, we propose a textual and structural dual enhancement (TSDE) framework that leverages large language models (LLMs) for graph data augmentation, thereby improving the performance of KGC models. To overcome the problem of low-quality entity descriptions, we designed a sampling algorithm to capture multiple paths surrounding entities, transforming this path information into text and integrating it into our carefully designed prompt templates to guide LLMs in generating high-quality descriptions enriched with contextual semantics. Additionally, to address the structural sparsity of KGs, we represent each entity description as a vector and calculate cosine similarity to identify a set of highly similar entities for each entity. We then assign similarity relations to these highly similar entity pairs to generate additional similar triples. After sampling these newly generated similar triples at a specified ratio, we incorporate them into the training set. We validated the effectiveness of the TSDE framework on two publicly available datasets: FB15k-237 and WN18RR. Experimental results demonstrate that the TSDE framework significantly enhances KGC model performance by refining entity descriptions and enriching KG structures.