ParaFusion-Extended: Large Scale Paraphrase Dataset Integrating Lexico-Phrasal Knowledge
摘要
Paraphrasing, the art of rephrasing text while retaining its original meaning, lies at the core of natural language understanding and generation. With the rise of demand for more domain-specialized models, high-quality data is more valued than ever; this includes paraphrasing. ParaFusion-Extend (PFE) is a large-scale dataset driven by Large Language Models incorporating lexical and phrasal knowledge. The dataset is curated to contain high-quality diverse paraphrase pairs and also separate knowledge bases that could be used for research work and data augmentation models. We show that PFE offers around at least a 30% increase in syntactic and lexical diversity compared to the original data sources that are commonly used. We demonstrate the effectiveness of PFE on several downstream tasks such as few-shot learning and training on sentence embeddings. We utilize a gold-standard evaluation scheme, which is further strengthened by human evaluation that shows the potential of PFE in advancing paraphrase generation.