<p>Knowledge graph question answering (KGQA) systems translate natural language questions into structured query languages (e.g., SQL/SPARQL). With advances in sequence-to-sequence models and large pre-trained language models (LPLMs), neural machine translation (NMT) has become a prevailing approach for Text-to-SPARQL. This study leverages the large-scale pre-trained language model <Emphasis FontCategory="NonProportional">T5</Emphasis> to obtain rich, transferable representations for SPARQL query generation and addresses translation errors that <Emphasis FontCategory="NonProportional">T5</Emphasis> may produce during decoding. Through an ablation study, we show that manually provided annotations (e.g., partial answers or gold entities) can effectively mitigate such errors; however, they require technical expertise and are therefore impractical for real-world deployment. To overcome this limitation, we propose a T5-based framework that integrates an <Emphasis FontCategory="NonProportional">MHC-LSTM</Emphasis> architecture with an <Emphasis FontCategory="NonProportional">automatic annotation and correction mechanism</Emphasis>. The automatic annotator, combined with the correction mechanism, yields the best overall results with <Emphasis FontCategory="NonProportional">T5-MHC-LSTM</Emphasis>, narrowing the gap to manually annotated performance. Empirically, the proposed method achieves F1-measures of <Emphasis FontCategory="NonProportional">89.63%</Emphasis> and <Emphasis FontCategory="NonProportional">95.58%</Emphasis> on <Emphasis FontCategory="NonProportional">QALD-9</Emphasis> and <Emphasis FontCategory="NonProportional">LC-QuAD 1.0</Emphasis> for text-to-text translation, and <Emphasis FontCategory="NonProportional">59%</Emphasis> and <Emphasis FontCategory="NonProportional">81%</Emphasis> in end-to-end evaluations, respectively—surpassing existing KGQA systems. These findings confirm that combining LPLMs, MHC-LSTM, and automated annotation with correction substantially enhances SPARQL query generation and overall KGQA effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic annotation for accurate text-to-SPARQL translation using hybrid encoder–decoder models

  • Yi-Hui Chen,
  • Eric Jui-Lin Lu,
  • Cheng-Hsien Hsu

摘要

Knowledge graph question answering (KGQA) systems translate natural language questions into structured query languages (e.g., SQL/SPARQL). With advances in sequence-to-sequence models and large pre-trained language models (LPLMs), neural machine translation (NMT) has become a prevailing approach for Text-to-SPARQL. This study leverages the large-scale pre-trained language model T5 to obtain rich, transferable representations for SPARQL query generation and addresses translation errors that T5 may produce during decoding. Through an ablation study, we show that manually provided annotations (e.g., partial answers or gold entities) can effectively mitigate such errors; however, they require technical expertise and are therefore impractical for real-world deployment. To overcome this limitation, we propose a T5-based framework that integrates an MHC-LSTM architecture with an automatic annotation and correction mechanism. The automatic annotator, combined with the correction mechanism, yields the best overall results with T5-MHC-LSTM, narrowing the gap to manually annotated performance. Empirically, the proposed method achieves F1-measures of 89.63% and 95.58% on QALD-9 and LC-QuAD 1.0 for text-to-text translation, and 59% and 81% in end-to-end evaluations, respectively—surpassing existing KGQA systems. These findings confirm that combining LPLMs, MHC-LSTM, and automated annotation with correction substantially enhances SPARQL query generation and overall KGQA effectiveness.