<p>Machine translation eliminates the obstacles caused by linguistic disparities around the world. The automatic translation of natural languages using machine translation methods breaks communication barriers and brings people closer together, regardless of language differences. Over the years, neural-based automatic natural language translation has achieved tremendous success. Despite its massive success, the neural-based approach is corpus-based, meaning that prediction accuracy depends on the input data volume. Recent years have witnessed an enormous growth in research into machine translation of various Indian languages. However, several Indian languages are still under investigation because their resources are inefficient for machine translation methods. In our experiment, we used a variety of neural-based techniques to assess the translation performance of the Nyishi-to-English corpora based on sentence lengths, a language with extremely limited online and offline resources. The experiment aims to identify the model’s performance based on the lengths of sentences with limited resources. We used the BLEU score up to 4-gram precision, Chrf, sacreBLEU and TER to judge the quality of prediction. Finally, we use the human evaluation method to evaluate the prediction error. We use BPE tokenization to handle the rare word difficulties of low-resource tonal languages. With BPE, all models with sentence lengths of 1–10 words and a BLEU score perform flawlessly.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An empirical investigation of the neural base approaches based on the sentence length using low-resource language: English-to-Nyishi

  • Nabam Kakum,
  • Rushanti Kri,
  • Koj Sambyo

摘要

Machine translation eliminates the obstacles caused by linguistic disparities around the world. The automatic translation of natural languages using machine translation methods breaks communication barriers and brings people closer together, regardless of language differences. Over the years, neural-based automatic natural language translation has achieved tremendous success. Despite its massive success, the neural-based approach is corpus-based, meaning that prediction accuracy depends on the input data volume. Recent years have witnessed an enormous growth in research into machine translation of various Indian languages. However, several Indian languages are still under investigation because their resources are inefficient for machine translation methods. In our experiment, we used a variety of neural-based techniques to assess the translation performance of the Nyishi-to-English corpora based on sentence lengths, a language with extremely limited online and offline resources. The experiment aims to identify the model’s performance based on the lengths of sentences with limited resources. We used the BLEU score up to 4-gram precision, Chrf, sacreBLEU and TER to judge the quality of prediction. Finally, we use the human evaluation method to evaluate the prediction error. We use BPE tokenization to handle the rare word difficulties of low-resource tonal languages. With BPE, all models with sentence lengths of 1–10 words and a BLEU score perform flawlessly.