An Experimental Research of Text-to-SQL for Heterogeneous Data in Large Language Models
摘要
The large language model (LLMs) technology has become a new paradigm for Text-to-SQL task. However, when faced with heterogeneous data, there is still a lack of comprehensive and effective solutions for Text-to-SQL task, especially in terms of high-quality corpus construction methods and low-resource training strategies. To address this challenge, this paper systematically studies Text-to-SQL technology in the field of heterogeneous data and innovatively proposes the HD-SQL method. This method achieves the construction of Text-to-SQL corpora for heterogeneous data that are secure, efficient, and low-cost. Based on this corpus, we further explore supervised fine-tuning strategies for low-resource LLMs on heterogeneous data. Compared to baseline models, significant performance improvements are achieved on the Heterogeneous Data Test set, further validating the practicality and effectiveness of the HD-SQL method. Overall, this paper proposes a feasible solution for Text-to-SQL task with heterogeneous data. We hope that our work can advance research in this field and provide strong technical support for practical applications.