Survey on the Availability of Datasets for Machine Translation Systems in Indian Languages
摘要
Machine translation is crucial to sharing knowledge using one language to any other language. It facilitates in communicating breaking language barrier. With the advent of neural networks, Machine translation has improved with good quality translations achieving human-like translation. Availability of high quality dataset remains a challenge. Since most of the Machine Translation work has been done on high resource languages like English and Chinese, very less dataset is available for Machine translation in Indian Languages. This paper tries to present an extensive list of existing dataset available for machine translation in Indian languages. These dataset are mostly pivoted to English language or not open accessible or include limited languages. We have discussed the importance of creating an open-source dataset that contains parallel sentences for multilingual translation models that includes underrepresented languages.