错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parallel Corpus Development Using Machine Translation

  • Sainik Kumar Mahata,
  • Jyoti Gupta,
  • Khusboo Kumari,
  • Monalisa Dey,
  • Anupam Mondal,
  • Darothi Sarkar

摘要

For a machine translation system to work efficiently, it requires training using a high-quality parallel corpus. But finding these types of parallel corpora is a problem for low-resourced languages. We plan to use machine translation methods to create a parallel corpus for English–Bengali language pairings in order to lessen this issue. For this, various translation systems were developed using the most recent methods and comparable corpora for English–Hindi was aligned using them. The alignment was based on automated metrics like BLEU and TER. After the corpus development, translation systems were trained using them and the quality was tested using the previous metrics and new manual metrics, Fluency and Adequacy. Testing showed that quality of translation produced using the later developed translation systems (using the developed parallel corpora) improved considerably.