错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging POS-Tag Features for Machine Translation of the Bengali–Nepali Language Pair: A Preliminary Study

  • Pooja Rai,
  • Sanjay Chatterji,
  • Samindra Basu

摘要

Machine Translation (MT) plays a crucial role in breaking language barriers and facilitating cross-lingual communication. However, certain language pairs, particularly those with limited linguistic resources, pose significant challenges in building effective MT systems. In this paper, we present the first attempt to build a machine translation system for the Bengali–Nepali language pair, which lacks an existing MT solution. To address the data scarcity issue, we utilized the Bengali–Nepali parallel treebanks, as the training resources for both Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) approaches. We adopted Moses, a widely used SMT framework, to develop an initial baseline system. Additionally, we integrated a factored language model that incorporates Part-of-Speech (POS) features to enhance the SMT-based translation performance. To further explore the potential of NMT for this low-resource language pair, we also experimented with transformer-based architecture. We leveraged the POS information to augment the transformer model and improve its translation capabilities. Our findings reveal that the inclusion of POS features in both SMT and NMT models leads to noticeable enhancements in translation quality. By building upon our initial findings, future research could potentially address the challenges posed by the scarcity of parallel data and contribute to more effective and reliable MT solutions for this specific language pair.