错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anticipating Retractions in Scientific Databases Using LLM-Based Citation Analysis

  • Muhammad Usman,
  • Mayukh Das,
  • Wolf-Tilo Balke

摘要

The increasing number of retracted articles raises concerns about scientific reliability, as they can spread flawed information. Moreover, uninformed and pre-retraction citations of these articles pose a cascading threat to scientific integrity. With no systematic method to identify articles at risk of retraction, we aim to address this challenge by detecting those susceptible to retraction and requiring further evaluation. We propose a triage process that focuses on Concerning Citations (CCs)—citations in which the citing paper questions the validity of a cited study’s data, methodology, or conclusions. To establish a foundation for this process, we developed a dedicated dataset to detect CCs, creating a new classification task for AI-based early identification of articles at risk of retraction. We evaluated machine learning (ML), encoder-based, and decoder-based large language models (LLMs) on this task, investigating the impact of scale, supervised fine-tuning (SFT), and few-shot learning on model performance. Our findings indicate that encoder-based models, particularly BERT, outperform other models. While scale and few-shot learning benefit large decoder models (e.g., LLaMA2), SFT does not consistently improve performance. This study contributes to the early identification of potential retractions, helping mitigate their impact on the scientific community.