Context and motivation. Large language models have the potential to bring important benefits in the context of processing high volumes of data and documents. LLMs led to the appearance of retrieval-augmented generation in which chunks of relevant data are being injected into large language models prompts in order to ensure the generation of higher quality content, anchored in verified facts and data. Question/problem. Large language model generation capabilities can be improved by providing parts of data in a more structured format, expressing information with a lesser amount of characters and tokens, without losing its semantic richness, situation where knowledge graphs can play a vital role in storing, at least partially, facts and data. Principal ideas/results. The paper at hand present a hybrid RAG approach in which data is gathered from both natural language documents and RDF knowledge graphs in order to provide an interplay between multiple heterogeneous data sources. Contribution. The presented hybrid RAG aims to improve the quality of LLM-based retrieval and querying of data, subsequently offering more flexibility when storing information.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Retrieval-Augmented Generation Approach for Heterogeneous Knowledge Bases

  • Andrei Chiș,
  • Ana-Maria Ghiran

摘要

Context and motivation. Large language models have the potential to bring important benefits in the context of processing high volumes of data and documents. LLMs led to the appearance of retrieval-augmented generation in which chunks of relevant data are being injected into large language models prompts in order to ensure the generation of higher quality content, anchored in verified facts and data. Question/problem. Large language model generation capabilities can be improved by providing parts of data in a more structured format, expressing information with a lesser amount of characters and tokens, without losing its semantic richness, situation where knowledge graphs can play a vital role in storing, at least partially, facts and data. Principal ideas/results. The paper at hand present a hybrid RAG approach in which data is gathered from both natural language documents and RDF knowledge graphs in order to provide an interplay between multiple heterogeneous data sources. Contribution. The presented hybrid RAG aims to improve the quality of LLM-based retrieval and querying of data, subsequently offering more flexibility when storing information.