<p>The powerful cross-linguistic capability of large language models has advanced the development of machine translation research. Based on the large language model of Xunzi’s series of ancient books, this paper explores the command fine-tuning of domain-oriented large models in machine translation oriented to the vertical domain of ancient books. One million two hundred thousand pairs of text-white parallel corpus were acquired, including three hundred thousand pairs of traditional corpus and nine hundred thousand pairs of simplified corpus, and the instruction fine-tuning dataset for the large language model was constructed. In this paper, six models are selected for instruction fine-tuning, and three different metrics are used to evaluate their performance. The evaluation results show that, compared with the generalized base model, the Xunzi series model improves in all indexes, and the Xunzi-Baichuan2-7B model has the best effect. Finally, the Xunzi-Baichuan2-7B model is fine-tuned with all parameters, and the translation results are analyzed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on machine translation of ancient books in the era of large language model

  • Zhixiao Zhao,
  • Guangyao Sun,
  • Chang Liu,
  • Dongbo Wang

摘要

The powerful cross-linguistic capability of large language models has advanced the development of machine translation research. Based on the large language model of Xunzi’s series of ancient books, this paper explores the command fine-tuning of domain-oriented large models in machine translation oriented to the vertical domain of ancient books. One million two hundred thousand pairs of text-white parallel corpus were acquired, including three hundred thousand pairs of traditional corpus and nine hundred thousand pairs of simplified corpus, and the instruction fine-tuning dataset for the large language model was constructed. In this paper, six models are selected for instruction fine-tuning, and three different metrics are used to evaluate their performance. The evaluation results show that, compared with the generalized base model, the Xunzi series model improves in all indexes, and the Xunzi-Baichuan2-7B model has the best effect. Finally, the Xunzi-Baichuan2-7B model is fine-tuned with all parameters, and the translation results are analyzed.