Enhancing low-resource neural machine translation with decoding-based data augmentation
摘要
Neural machine translation (NMT) for low-resource languages remains challenging due to the limited availability of high-quality parallel corpora. This paper proposes a novel data augmentation strategy that integrates beam search, restricted sampling, and paraphrasing to improve translation quality without relying on additional monolingual data. By generating diverse and representative pseudo-parallel sentence pairs from existing training data, the proposed method enhances generalization and robustness of NMT models. The framework is evaluated on the Kashmiri-English language pair and further tested on a secondary language pair to demonstrate its adaptability. Results confirm that the approach effectively addresses data scarcity while maintaining computational efficiency, making it suitable for broader application in low-resource settings.