Parallel Corpus Development Using Machine Translation
摘要
For a machine translation system to work efficiently, it requires training using a high-quality parallel corpus. But finding these types of parallel corpora is a problem for low-resourced languages. We plan to use machine translation methods to create a parallel corpus for English–Bengali language pairings in order to lessen this issue. For this, various translation systems were developed using the most recent methods and comparable corpora for English–Hindi was aligned using them. The alignment was based on automated metrics like BLEU and TER. After the corpus development, translation systems were trained using them and the quality was tested using the previous metrics and new manual metrics, Fluency and Adequacy. Testing showed that quality of translation produced using the later developed translation systems (using the developed parallel corpora) improved considerably.