<p>Conversational services based on large language models emerge as a promising service delivery paradigm, which are notably characterized by intuitive natural language interaction and highly efficient service fulfillment. However, their adoption in government services encounters distinctive challenges due to the stringent requirements of the domain for precise user demand satisfaction, which mandates models retrieve massive cross-domain external documents with complex multi-hop dependencies. When faced with such multi-document question answering (MD-QA) problems, the most popular approach is to initially utilize multi-hop dense retrieval (MDR) for pre-constructing document graphs, and subsequently traverse the document graphs with the Knowledge Graph Prompting (KGP) method. However, the graphs constructed by this method hardly reach the expected quality level, and the phenomenon of hallucination also occurs occasionally, which is unacceptable for government services. To address these issues, we make the following improvements. Firstly, we propose the <Emphasis Type="BoldItalic">SubMDR</Emphasis> model, which adopts a proposition-based data augmentation strategy to optimize MDR training and construct higher-quality graphs. Secondly, we introduce <Emphasis Type="BoldItalic">KGP3</Emphasis> (Knowledge Graph Prompting with a Three-Stage Traversal Framework), which improves graph traversal and reasoning by combining sub-question generation, neighbor-node selection, and node verification. Finally, we construct an augmented dataset to finetune models of different scales. Experiments on HotpotQA and MuSiQue demonstrate significant improvements, enabling more reliable service delivery.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Multi-Document Question Answering with Semantic Document Graph Construction and LLM-Based Graph Retrieval

  • Zedong Jin,
  • Zhiwei Yi,
  • Zhiying Tu,
  • Dianhui Chu

摘要

Conversational services based on large language models emerge as a promising service delivery paradigm, which are notably characterized by intuitive natural language interaction and highly efficient service fulfillment. However, their adoption in government services encounters distinctive challenges due to the stringent requirements of the domain for precise user demand satisfaction, which mandates models retrieve massive cross-domain external documents with complex multi-hop dependencies. When faced with such multi-document question answering (MD-QA) problems, the most popular approach is to initially utilize multi-hop dense retrieval (MDR) for pre-constructing document graphs, and subsequently traverse the document graphs with the Knowledge Graph Prompting (KGP) method. However, the graphs constructed by this method hardly reach the expected quality level, and the phenomenon of hallucination also occurs occasionally, which is unacceptable for government services. To address these issues, we make the following improvements. Firstly, we propose the SubMDR model, which adopts a proposition-based data augmentation strategy to optimize MDR training and construct higher-quality graphs. Secondly, we introduce KGP3 (Knowledge Graph Prompting with a Three-Stage Traversal Framework), which improves graph traversal and reasoning by combining sub-question generation, neighbor-node selection, and node verification. Finally, we construct an augmented dataset to finetune models of different scales. Experiments on HotpotQA and MuSiQue demonstrate significant improvements, enabling more reliable service delivery.