Improving Quality of Vietnamese to Khmer Neural Machine Translation Using Multi-stage Fine-Tuning Strategy
摘要
Machine translation for low-resource language pairs remains a challenging task, even with the advanced capabilities of large language models. Multilingual large language models require extensive monolingual data during the pre-training phase and large, high-quality parallel datasets for fine-tuning. In this paper, we present our research on an effective fine-tuning strategy for a pre-trained large language model to improve machine translation quality for the Vietnamese-Khmer language pair. Our experiments show that by applying self-supervised learning to the pre-trained model, fine-tuning it on related tasks, and then further fine-tuning it on both the original and augmented datasets, we achieved a BLEU score improvement of over 13% compared to the best results from previous studies and 7% higher compared to Google Translator and GPT-4 for in-domain tests.