Low-Resource Indic Languages Translation Using Multilingual Approaches
摘要
Machine translation is effective in the presence of a substantial parallel corpus. In a multilingual country like India, with diverse linguistic origins and scripts, the vast majority of languages need more resources to produce high-quality translation models. Multilingual neural machine translation (MNMT) has the advantage of being scalable across multiple languages and improving low-resource languages via knowledge transfer. In this work, we investigate MNMT for low-resource Indic languages—Hindi, Bengali, Assamese, Manipuri, and Mizo. With the recent success of massively multilingual pre-trained models for low-resource languages, we explore the effectiveness of using multilingual pre-trained transformers—mBART and mT5 on several Indic languages. We perform fine-tuning on the pre-trained models in a one-to-many and many-to-one approach. We compare the performance of multilingual pre-trained models with multiway multilingual translation trained from scratch using a one-to-many and many-to-one approach.