Abstract <p>The exponential growth of scientific publications has heightened the need for robust tools to organize and retrieve research effectively. The Universal Decimal Classification (UDC) serves as a valuable framework for categorizing articles by subject area. However, manual assignment of UDC codes is often prone to inaccuracies or oversimplification, limiting its utility. In this study, we present a novel approach for the automated assignment of UDC codes to scientific articles using BERT-based models. Our methodology is trained and evaluated on a dataset comprising over 19&#xa0;000 articles in mathematics and related disciplines. To address the hierarchical structure of the UDC, we develop two specialized evaluation metrics: hierarchical classification accuracy and hierarchical recommendation accuracy. We also explore multiple strategies for flattening hierarchical labels. Our results demonstrate a hierarchical recommendation accuracy of 0.8220. Furthermore, blind expert evaluation reveals that discrepancies between the reference and predicted labels often stem from errors in the original UDC code assignments by the authors of articles. Our approach demonstrates strong potential for automating the classification of scientific articles and can be extended to other hierarchical classification systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Classification of Scientific Articles Using Deep Learning (Using the UDC Hierarchy As an Example)

  • V. Y. Mamedov,
  • D. A. Kovalevsky,
  • D. A. Morozov,
  • S. S. Stolyarov,
  • S. S. Ospichev

摘要

Abstract

The exponential growth of scientific publications has heightened the need for robust tools to organize and retrieve research effectively. The Universal Decimal Classification (UDC) serves as a valuable framework for categorizing articles by subject area. However, manual assignment of UDC codes is often prone to inaccuracies or oversimplification, limiting its utility. In this study, we present a novel approach for the automated assignment of UDC codes to scientific articles using BERT-based models. Our methodology is trained and evaluated on a dataset comprising over 19 000 articles in mathematics and related disciplines. To address the hierarchical structure of the UDC, we develop two specialized evaluation metrics: hierarchical classification accuracy and hierarchical recommendation accuracy. We also explore multiple strategies for flattening hierarchical labels. Our results demonstrate a hierarchical recommendation accuracy of 0.8220. Furthermore, blind expert evaluation reveals that discrepancies between the reference and predicted labels often stem from errors in the original UDC code assignments by the authors of articles. Our approach demonstrates strong potential for automating the classification of scientific articles and can be extended to other hierarchical classification systems.