The research focuses on citation classification as influential or non-influential by analyzing the importance of data driven factors. Citation analysis is crucial for measuring the impact of academic work and identifying emerging research topics. However, not all citations hold equal importance. Previous citation classification studies use two frameworks: metadata analysis (e.g., title, author, venue) and contextual analysis, which assesses the impact based on how citations appear in the citing work’s body. In this research the use of diverse features (textual, cue word-based, contextual) to measure impact factors has been highlighted. The significance of these features in different scenarios also has been illustrated. An automatic classification mechanism combining qualitative and quantitative analysis has been provided. Four classification techniques (Random Forest, Decision Tree, Support Vector Machine, and Neural Network) on an annotated dataset of 412 citations from 20,527 papers in the Association for Computational Linguistics Anthology have been applied. Analyzing 23 unique features, ROC curves are used to evaluate performance. Random Forest and Support Vector Machine achieved the highest AUC scores of 0.929 and 0.927, respectively, outperforming Neural Network and Decision Tree, which scored 0.893 and 0.906. The research approach surpassed baseline methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Techniques to Identify the Importance of Data Driven Factors in Citation Classification

  • Waseem Bacha,
  • Arifa Ashrafi,
  • Victor Sergeevich Mokhnachev,
  • Ali Daud

摘要

The research focuses on citation classification as influential or non-influential by analyzing the importance of data driven factors. Citation analysis is crucial for measuring the impact of academic work and identifying emerging research topics. However, not all citations hold equal importance. Previous citation classification studies use two frameworks: metadata analysis (e.g., title, author, venue) and contextual analysis, which assesses the impact based on how citations appear in the citing work’s body. In this research the use of diverse features (textual, cue word-based, contextual) to measure impact factors has been highlighted. The significance of these features in different scenarios also has been illustrated. An automatic classification mechanism combining qualitative and quantitative analysis has been provided. Four classification techniques (Random Forest, Decision Tree, Support Vector Machine, and Neural Network) on an annotated dataset of 412 citations from 20,527 papers in the Association for Computational Linguistics Anthology have been applied. Analyzing 23 unique features, ROC curves are used to evaluate performance. Random Forest and Support Vector Machine achieved the highest AUC scores of 0.929 and 0.927, respectively, outperforming Neural Network and Decision Tree, which scored 0.893 and 0.906. The research approach surpassed baseline methods.