Large language models (LLMs) are increasingly used for table+text question answering (QA), but there is insufficient focus on unified table and text representation needed for effective retrieval guidance. In response, we present TabSegNet, comprised of (1) a fine-tuned LLM to decompose the question into segments that access individual tables;(2) a unified graph representation of tables+text, where retrieval amounts to identifying a compact node subset collectively covering the question segments; and (3) a novel graph neural network (GNN) whose messages are informed by the question and its segmentation. Experiments with existing and newly-created data sets demonstrate the promise of our approach, compared to sparse and dense nearest neighbor search, or using an LLM for retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph Representation of Tables+Text and Compact Subgraph Retrieval for QA Tasks

  • Vishwajeet Kumar,
  • Jaydeep Sen,
  • Bhawna Chelani,
  • Soumen Chakrabarti

摘要

Large language models (LLMs) are increasingly used for table+text question answering (QA), but there is insufficient focus on unified table and text representation needed for effective retrieval guidance. In response, we present TabSegNet, comprised of (1) a fine-tuned LLM to decompose the question into segments that access individual tables;(2) a unified graph representation of tables+text, where retrieval amounts to identifying a compact node subset collectively covering the question segments; and (3) a novel graph neural network (GNN) whose messages are informed by the question and its segmentation. Experiments with existing and newly-created data sets demonstrate the promise of our approach, compared to sparse and dense nearest neighbor search, or using an LLM for retrieval.