Scaling Up Paraphrase Generation Datasets with Machine Translation and Semantic Similarity Filtering
摘要
Paraphrase generation is an important natural language processing task that can be used in other tasks such as machine translation, question answering, and semantic parsing. The problem of data scarcity remains the main obstacle in the way of developing models for this task, especially for low-resource languages. In this chapter, we introduce large paraphrase corpora and a new methodology for filtering non-relevant sentence pairs. We augmented some datasets using a model, we fine-tuned on a dataset created using our method, and we obtained satisfactory results asserting our data collection and filtering method’s effectiveness.