An Enhanced Method for Mongolian-Chinese Neural Machine Translation Using Multilingual Datastores and Chinese-Centric Methods
摘要
Low-resource language translation remains a significant challenge in natural language processing, particularly for the Mongolian-Chinese language pair under the “Belt and Road” initiative. Existing translation systems struggle with this pair due to the scarcity of high-quality data. This paper addresses these challenges by combining multilingual k-nearest-neighbor machine translation (kNN-MT) with Chinese-centric methods. We constructed a robust multilingual datastore and introduced an incomplete-trust loss function to effectively manage low-quality data. Additionally, we implemented re-ranking techniques to further enhance the robustness and accuracy of the translation model. The experimental results indicate that this combined approach significantly improves Mongolian-Chinese translation quality on the mBART model, with a BLEU score increase of 3.81 points and a TER score decrease of 0.0531 points. Our findings demonstrate that integrating kNN-MT with Chinese-centric methods and employing advanced loss functions and re-ranking techniques can effectively address data scarcity and quality issues, leading to substantial improvements in translation performance for low-resource language pairs.