Efficient organization and access to scholarly information are paramount in the digital age, where the volume of research papers is ever-growing. Research paper classification is a crucial task that aids in organizing and retrieving this vast amount of information. This study leverages the DistilBERT model, a distilled version of the BERT (Bidirectional Encoder Representations from Transformers) model, to classify research papers based on their titles and summaries. DistilBERT offers a computationally efficient solution for natural language processing tasks, making it ideal for large-scale document classification. The proposed method involves fine-tuning DistilBERT on a dataset of research papers to develop a classification model. Initially, the titles and summaries of research papers are preprocessed to convert them into vectors using DistilBERT’s pre-trained language model. These vectors are then used as input features to train a classifier that predicts the category of the research paper. The proposed approach was evaluated on a dataset of research papers across various disciplines. The results demonstrate the method’s effectiveness, with classification performance metrics such as F1 score, recall, precision, and accuracy all exceeding 81%. These results indicate that the proposed method can accurately classify research papers while significantly reducing the computational resources required compared to the original BERT model. Efficient classification of research papers can greatly benefit researchers, educators, and students by providing them with a more organized and accessible repository of scholarly information. The study highlights the utility of transformer models and transfer learning in scholarly article classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scholarly Article Classification Leveraging DistilBERT Transformer and Transfer Learning

  • Rasha S. Gargees

摘要

Efficient organization and access to scholarly information are paramount in the digital age, where the volume of research papers is ever-growing. Research paper classification is a crucial task that aids in organizing and retrieving this vast amount of information. This study leverages the DistilBERT model, a distilled version of the BERT (Bidirectional Encoder Representations from Transformers) model, to classify research papers based on their titles and summaries. DistilBERT offers a computationally efficient solution for natural language processing tasks, making it ideal for large-scale document classification. The proposed method involves fine-tuning DistilBERT on a dataset of research papers to develop a classification model. Initially, the titles and summaries of research papers are preprocessed to convert them into vectors using DistilBERT’s pre-trained language model. These vectors are then used as input features to train a classifier that predicts the category of the research paper. The proposed approach was evaluated on a dataset of research papers across various disciplines. The results demonstrate the method’s effectiveness, with classification performance metrics such as F1 score, recall, precision, and accuracy all exceeding 81%. These results indicate that the proposed method can accurately classify research papers while significantly reducing the computational resources required compared to the original BERT model. Efficient classification of research papers can greatly benefit researchers, educators, and students by providing them with a more organized and accessible repository of scholarly information. The study highlights the utility of transformer models and transfer learning in scholarly article classification.