Low-resource language translation remains a significant challenge in natural language processing, particularly for the Mongolian-Chinese language pair under the “Belt and Road” initiative. Existing translation systems struggle with this pair due to the scarcity of high-quality data. This paper addresses these challenges by combining multilingual k-nearest-neighbor machine translation (kNN-MT) with Chinese-centric methods. We constructed a robust multilingual datastore and introduced an incomplete-trust loss function to effectively manage low-quality data. Additionally, we implemented re-ranking techniques to further enhance the robustness and accuracy of the translation model. The experimental results indicate that this combined approach significantly improves Mongolian-Chinese translation quality on the mBART model, with a BLEU score increase of 3.81 points and a TER score decrease of 0.0531 points. Our findings demonstrate that integrating kNN-MT with Chinese-centric methods and employing advanced loss functions and re-ranking techniques can effectively address data scarcity and quality issues, leading to substantial improvements in translation performance for low-resource language pairs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Enhanced Method for Mongolian-Chinese Neural Machine Translation Using Multilingual Datastores and Chinese-Centric Methods

  • Bailun Wang,
  • Yatu Ji,
  • Nier Wu,
  • Xu Liu,
  • Yanli Wang,
  • Rui Mao,
  • Chao Zhou,
  • Yepai Jia,
  • Chen Zhao,
  • Qing-Dao-Er-Ji Ren,
  • Na Liu

摘要

Low-resource language translation remains a significant challenge in natural language processing, particularly for the Mongolian-Chinese language pair under the “Belt and Road” initiative. Existing translation systems struggle with this pair due to the scarcity of high-quality data. This paper addresses these challenges by combining multilingual k-nearest-neighbor machine translation (kNN-MT) with Chinese-centric methods. We constructed a robust multilingual datastore and introduced an incomplete-trust loss function to effectively manage low-quality data. Additionally, we implemented re-ranking techniques to further enhance the robustness and accuracy of the translation model. The experimental results indicate that this combined approach significantly improves Mongolian-Chinese translation quality on the mBART model, with a BLEU score increase of 3.81 points and a TER score decrease of 0.0531 points. Our findings demonstrate that integrating kNN-MT with Chinese-centric methods and employing advanced loss functions and re-ranking techniques can effectively address data scarcity and quality issues, leading to substantial improvements in translation performance for low-resource language pairs.