<p>Neural machine translation (NMT) for low-resource languages remains challenging due to the limited availability of high-quality parallel corpora. This paper proposes a novel data augmentation strategy that integrates beam search, restricted sampling, and paraphrasing to improve translation quality without relying on additional monolingual data. By generating diverse and representative pseudo-parallel sentence pairs from existing training data, the proposed method enhances generalization and robustness of NMT models. The framework is evaluated on the Kashmiri-English language pair and further tested on a secondary language pair to demonstrate its adaptability. Results confirm that the approach effectively addresses data scarcity while maintaining computational efficiency, making it suitable for broader application in low-resource settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing low-resource neural machine translation with decoding-based data augmentation

  • Syed Matla Ul Qumar,
  • Muzaffar Azim,
  • S. M. K. Quadri

摘要

Neural machine translation (NMT) for low-resource languages remains challenging due to the limited availability of high-quality parallel corpora. This paper proposes a novel data augmentation strategy that integrates beam search, restricted sampling, and paraphrasing to improve translation quality without relying on additional monolingual data. By generating diverse and representative pseudo-parallel sentence pairs from existing training data, the proposed method enhances generalization and robustness of NMT models. The framework is evaluated on the Kashmiri-English language pair and further tested on a secondary language pair to demonstrate its adaptability. Results confirm that the approach effectively addresses data scarcity while maintaining computational efficiency, making it suitable for broader application in low-resource settings.