Collection and Preprocessing of Data for LLM in the Kazakh Language in the Field of Legislation
摘要
This article presents the process of preparing a dataset for a question-answer system in the state language in the field of legislation of the Republic of Kazakhstan. The data were web crawled from the websites of the Information and Legal System of Normative Legal Acts of the Republic of Kazakhstan “Adilet” and the “Institute of Legislation and Legal Information of the Republic of Kazakhstan” of the Ministry of Justice of the Republic of Kazakhstan (Zqai). The collection of these datasets took a rigorous parsing process to ensure the achievement of data quality and consistency. The first formed dataset consists of the header, date, text, and source fields, while the second one includes the number, date, question, and answer features. The amount of data collected from the ‘Adilet’ website is more than 500 thousand sentences, and the question-answer dataset from the ‘Zqai’ website consists of 740 questions and answers with a size of more than 18 thousand sentences. These datasets have a gigantic meaning in large language models where the construction of advanced question-answer systems is of great meaning. All of this will allow assisting users with legal queries in natural languages.