错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data-Efficient Training for Effective Paraphrase Retrieval Techniques Using Language Models to Identify Research Gaps

  • Deepa Shree Chickballapur Venkatachalapathi,
  • Chandrashekhar Pomu Chavan

摘要

Efficiently identifying research gaps is a critical aspect of scholarly endeavors, often demanding extensive literature surveys that require huge amounts of manual effort. Enhancing this process using Machine Learning techniques can be time consuming and expensive, as huge amounts of data needs to be used to come up with high accuracy models. This paper proposes an approach to streamline this process by integrating semantic search with paraphrase retrieval and employing few-shot learning techniques with language models. The primary objective is to perform paraphrase retrieval at the sentence level using data-efficient techniques, thereby reducing the latency and cost associated with model training. In this study, multiple models were fine-tuned for paraphrase detection, and the most efficient among them was employed for paraphrase retrieval on academic works during testing. Libraries such as SetFit and few-shot learning techniques were applied with minimal training sample sizes, yielding accuracies exceeding 90%. The model was then evaluated to identify paraphrase pairs within text, enhancing the precision of search results with fewer queries. The paper also discusses the limitations of the proposed approaches and outlines plans for further enhancements. The novelty of this research lies in the combination of paraphrase detection and few-shot learning techniques for semantic searching, offering a more efficient and effective approach to identify research gaps.