An empirical investigation of the neural base approaches based on the sentence length using low-resource language: English-to-Nyishi
摘要
Machine translation eliminates the obstacles caused by linguistic disparities around the world. The automatic translation of natural languages using machine translation methods breaks communication barriers and brings people closer together, regardless of language differences. Over the years, neural-based automatic natural language translation has achieved tremendous success. Despite its massive success, the neural-based approach is corpus-based, meaning that prediction accuracy depends on the input data volume. Recent years have witnessed an enormous growth in research into machine translation of various Indian languages. However, several Indian languages are still under investigation because their resources are inefficient for machine translation methods. In our experiment, we used a variety of neural-based techniques to assess the translation performance of the Nyishi-to-English corpora based on sentence lengths, a language with extremely limited online and offline resources. The experiment aims to identify the model’s performance based on the lengths of sentences with limited resources. We used the BLEU score up to 4-gram precision, Chrf, sacreBLEU and TER to judge the quality of prediction. Finally, we use the human evaluation method to evaluate the prediction error. We use BPE tokenization to handle the rare word difficulties of low-resource tonal languages. With BPE, all models with sentence lengths of 1–10 words and a BLEU score perform flawlessly.