The dazzling brilliance of deep learning makes it a consensus that “give me large-scale high-quality sentence pairs to train on, I can obtain a high-performance neural machine translation model.” The current ever-increasing productivity of human language data greatly reduces the scarcity of sentence pairs, but the explosion of language data also brings potential problems such as uneven translation quality of sentences. We address the issue of sentence pair filtering based on translation quality, explore a new method of morphological semantic ensemble filtering, which can fully leverage the efficiency of morphological measurement derived from Levenshtein editing distance and the accuracy of semantic measurement derived from pretrained models, and achieve an efficient estimation of sentence pair translation quality taking into account both morphology and semantics. We first conduct morphological filtering, semantic filtering, and morphological semantic ensemble filtering experiments on the datasets of 17 language pairs respectively, and then use the filtered sentence pairs to enhance the retraining of the 17 neural machine translation models. Experimental results show that all three filtered results can significantly improve the performance of neural machine translation models, among which, the morphological semantic ensemble filtering has the best effect, improving 2–3 BLEU points compared with the other two methods. Experimental results clarify that our new method can use morphological computing to strengthen perfect homonym data, while can use semantic computing to strengthen shaped synonym data. The corresponding implementation is an industrial-level straightforward and efficient algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Morphological Semantic Ensemble Filtering of Massive Sentence Pairs for Neural Machine Translation

  • Lin Wang,
  • Wuying Liu

摘要

The dazzling brilliance of deep learning makes it a consensus that “give me large-scale high-quality sentence pairs to train on, I can obtain a high-performance neural machine translation model.” The current ever-increasing productivity of human language data greatly reduces the scarcity of sentence pairs, but the explosion of language data also brings potential problems such as uneven translation quality of sentences. We address the issue of sentence pair filtering based on translation quality, explore a new method of morphological semantic ensemble filtering, which can fully leverage the efficiency of morphological measurement derived from Levenshtein editing distance and the accuracy of semantic measurement derived from pretrained models, and achieve an efficient estimation of sentence pair translation quality taking into account both morphology and semantics. We first conduct morphological filtering, semantic filtering, and morphological semantic ensemble filtering experiments on the datasets of 17 language pairs respectively, and then use the filtered sentence pairs to enhance the retraining of the 17 neural machine translation models. Experimental results show that all three filtered results can significantly improve the performance of neural machine translation models, among which, the morphological semantic ensemble filtering has the best effect, improving 2–3 BLEU points compared with the other two methods. Experimental results clarify that our new method can use morphological computing to strengthen perfect homonym data, while can use semantic computing to strengthen shaped synonym data. The corresponding implementation is an industrial-level straightforward and efficient algorithm.