Accurately evaluating research impact is crucial in academia, influencing funding, promotions, and recognition. However, exces sive self-citations distort citation metrics, undermining fair assessment. This work introduces a novel approach to detect anomalous self-citations using citation network analysis and advanced Natural Language Processing with Large Language Models. A citation network is constructed from a large-scale academic dataset, where nodes represent papers and authors, and edges capture citation relationships. Self-citation loops are identified using graph-based techniques. A two-stage summarization process is implemented to generate a comprehensive summary. Regular expressions facilitate citation context detection, while prompt fine-tuning through self-contrast improves the LLMs’ ability to classify essential vs. non-essential citations, reducing prompt loss to 0.082. Extensive testing confirms the effectiveness of this approach, with o1-mini achieving 91.84% accuracy on 49 self-citation cases across a set of authors. The findings provide actionable insights to enhance transparency and fairness in research evaluation. By addressing ethical concerns in scholarly publishing, this research promotes integrity and equitable academic assessments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Anomalous Self-citations Using Citation Network Analysis and LLMs

  • Farhaan Ebadulla,
  • B. V. Gaurav,
  • H. Manoj,
  • Arti Arya,
  • Nazmin Begum,
  • K. Kruthik

摘要

Accurately evaluating research impact is crucial in academia, influencing funding, promotions, and recognition. However, exces sive self-citations distort citation metrics, undermining fair assessment. This work introduces a novel approach to detect anomalous self-citations using citation network analysis and advanced Natural Language Processing with Large Language Models. A citation network is constructed from a large-scale academic dataset, where nodes represent papers and authors, and edges capture citation relationships. Self-citation loops are identified using graph-based techniques. A two-stage summarization process is implemented to generate a comprehensive summary. Regular expressions facilitate citation context detection, while prompt fine-tuning through self-contrast improves the LLMs’ ability to classify essential vs. non-essential citations, reducing prompt loss to 0.082. Extensive testing confirms the effectiveness of this approach, with o1-mini achieving 91.84% accuracy on 49 self-citation cases across a set of authors. The findings provide actionable insights to enhance transparency and fairness in research evaluation. By addressing ethical concerns in scholarly publishing, this research promotes integrity and equitable academic assessments.