Knowledge Graphs (KGs) have been broadly adopted as representation models to capture the world’s information in a flexible way. This has allowed them to be the backbone of many different information systems, ranging from Web search engines to integration solutions for in-house enterprise systems. Given that they contain knowledge about entities, we need mechanisms to efficiently search for them, which leads to the task of Entity Retrieval. Searching for particular entities in a KG is usually tackled by applying traditional Information Retrieval (IR) techniques: indexing the KG by building Virtual Documents for each entity and applying classic IR models. However, this requires a deep analysis of the ontology of the KG (e.g., its properties and types) and the distribution of the actual data populating it. Manually analyzing these aspects in large KGs to build custom-tailored indexes becomes unfeasible, requiring instead supervised Machine Learning approaches. In this work, we propose an approach to creating a multi-field and information-aware KG index in a completely unsupervised way. By building on different information measures, our approach clusters the properties of the KG by capturing their relative importance at different levels. This allows us to build a retrieval system for all the entities in the KG, or focus on a subset of them based on their types (i.e., building a vertical search engine). Experimental results on the DBpedia and IMDb KGs show that our retrieval performance rivals that of the state of the art, without requiring either: 1) any manual analysis of the graph regardless of its size, type hierarchy or general complexity of its properties, or 2) the use of any supervised Machine Learning techniques in order to fine-tune fielded retrieval models. Thus, we additionally eliminate the need to create annotated data such as relevance query sets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Information-Aware Entity Indexing in Knowledge Graphs to Enable Semantic Search

  • Samuel García,
  • Carlos Bobed

摘要

Knowledge Graphs (KGs) have been broadly adopted as representation models to capture the world’s information in a flexible way. This has allowed them to be the backbone of many different information systems, ranging from Web search engines to integration solutions for in-house enterprise systems. Given that they contain knowledge about entities, we need mechanisms to efficiently search for them, which leads to the task of Entity Retrieval. Searching for particular entities in a KG is usually tackled by applying traditional Information Retrieval (IR) techniques: indexing the KG by building Virtual Documents for each entity and applying classic IR models. However, this requires a deep analysis of the ontology of the KG (e.g., its properties and types) and the distribution of the actual data populating it. Manually analyzing these aspects in large KGs to build custom-tailored indexes becomes unfeasible, requiring instead supervised Machine Learning approaches. In this work, we propose an approach to creating a multi-field and information-aware KG index in a completely unsupervised way. By building on different information measures, our approach clusters the properties of the KG by capturing their relative importance at different levels. This allows us to build a retrieval system for all the entities in the KG, or focus on a subset of them based on their types (i.e., building a vertical search engine). Experimental results on the DBpedia and IMDb KGs show that our retrieval performance rivals that of the state of the art, without requiring either: 1) any manual analysis of the graph regardless of its size, type hierarchy or general complexity of its properties, or 2) the use of any supervised Machine Learning techniques in order to fine-tune fielded retrieval models. Thus, we additionally eliminate the need to create annotated data such as relevance query sets.