错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving the Accuracy of Text-to-SQL Tools Based on Large Language Models for Real-World Relational Databases

  • Gustavo M. C. Coelho,
  • Eduardo R. S. Nascimento,
  • Yenier T. Izquierdo,
  • Grettel M. García,
  • Lucas Feijó,
  • Melissa Lemos,
  • Robinson L. S. Garcia,
  • Aiko R. de Oliveira,
  • João P. Pinheiro,
  • Marco A. Casanova

摘要

Real-world relational databases (RW-RDB) have large, complex schemas often expressed in terms alien to end-users. This scenario is challenging to LLM-based text-to-SQL tools, that is, tools that translate Natural Language (NL) sentences into SQL queries using a Large Language Model (LLM). Indeed, their accuracy on RW-RDBs is considerably less than that reported for well-known synthetic benchmarks. This paper then introduces a technique to improve the accuracy of LLM-based text-to-SQL tools on RW-RDBs using Retrieval-Augmented Generation. The technique consists of two steps. Using the RW-RDB schema, the first step generates a synthetic dataset E of pairs \((Q_N,Q_S)\) , where \(Q_N\) is an NL sentence and \(Q_S\) is the corresponding SQL translation. The core contribution of the paper is an algorithm that implements this first step. Given an input NL sentence \(Q_I\) , the second step retrieves pairs \((Q_N,Q_S)\) from E based on the similarity of \(Q_I\) and \(Q_N\) , and prompts such pairs to the LLM to improve accuracy. To argue in favor of the proposed technique, the paper includes experiments with an RW-RDB, which is in production at an Energy company, and a well-known text-to-SQL prompt strategy. It repeats the experiments with Mondial, an openly available database with a large schema. These experiments constitute a second contribution of the paper.