A Source Template-Based Data Augmentation Method for Low-Resource Neural Machine Translation
摘要
Neural machine translation (NMT) has recently garnered considerable attention, mainly because of its inherent ability to produce highly precise translations. However, the effectiveness of NMT model heavily relies on the availability of extensive training data, and the performance of the translation model tends to degrade markedly in the absence of large-scale, high-quality datasets. To alleviate this challenge, especially for low-resource languages, a data augmentation (DA) approach based on source templates is proposed. Initially, a template extraction algorithm is introduced, which is applied to each source sentence in the training corpus. Subsequently, pseudo-parallel data re constructed by pairing the generated source templates with their corresponding original target sentences. Lastly, two DA strategies are employed to enrich the training set. Experimental results demonstrate that, in comparison to both the baseline model and several existing DA methods, the proposed source template-based DA method effectively enhances translation quality.