Currently, data augmentation is a primary technique for improving the performance of Neural Machine Translation (NMT) in low-resource settings. However, traditional data augmentation methods, while alleviating issues related to data sparsity, are often influenced by noise, leading to syntactic errors and semantic discrepancies, which in turn degrade the quality of NMT outputs. To address this challenge, this paper proposes a self-supervised data augmentation approach that integrates contrastive learning to pull similar features closer together, thereby reducing the noise introduced by conventional augmentation techniques. Moreover, to resolve the issue of excessive low-frequency words in traditional low-resource agglutinative language NMT, commonly used solutions such as stem-and-affix segmentation can preserve basic semantic information but depend on manually curated dictionaries, which lack flexibility. Although Byte Pair Encoding (BPE) is a statistical method, it fails to capture word-level semantic features. In light of these challenges, this paper introduces a morphological recombination approach to further enhance translation quality. Specifically, we propose a convolutional gated morphological attention mechanism in the encoder to capture and amplify morphological features within the input sequence, while a morphological cross-attention mechanism in the decoder ensures that these features are effectively leveraged to guide the translation process. Experiments conducted on the Mn-Zh task using various gated attention mechanisms demonstrate an average BLEU score improvement of 2.62%, validating the efficacy of the proposed model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Morphological Recombination-Based Neural Machine Translation with Self-Supervised Data Augmentation

  • Yukun He,
  • Nier Wu,
  • Yuxin Zhao,
  • Menghan Li

摘要

Currently, data augmentation is a primary technique for improving the performance of Neural Machine Translation (NMT) in low-resource settings. However, traditional data augmentation methods, while alleviating issues related to data sparsity, are often influenced by noise, leading to syntactic errors and semantic discrepancies, which in turn degrade the quality of NMT outputs. To address this challenge, this paper proposes a self-supervised data augmentation approach that integrates contrastive learning to pull similar features closer together, thereby reducing the noise introduced by conventional augmentation techniques. Moreover, to resolve the issue of excessive low-frequency words in traditional low-resource agglutinative language NMT, commonly used solutions such as stem-and-affix segmentation can preserve basic semantic information but depend on manually curated dictionaries, which lack flexibility. Although Byte Pair Encoding (BPE) is a statistical method, it fails to capture word-level semantic features. In light of these challenges, this paper introduces a morphological recombination approach to further enhance translation quality. Specifically, we propose a convolutional gated morphological attention mechanism in the encoder to capture and amplify morphological features within the input sequence, while a morphological cross-attention mechanism in the decoder ensures that these features are effectively leveraged to guide the translation process. Experiments conducted on the Mn-Zh task using various gated attention mechanisms demonstrate an average BLEU score improvement of 2.62%, validating the efficacy of the proposed model.