Enriching Ontology with Named Entity Recognition (NER) Integration
摘要
Our work focuses on the development of an innovative search and annotation system for arXiv, a platform renowned for its advanced capabilities in exploring research articles, with a particular emphasis on machine learning. Our approach, which enhances content through contextual annotations, goes beyond traditional categorizations. This involves a thorough exploration of machine learning works on arXiv, complementing the existing search features of the platform such as categories, authors, dates, affiliations, etc. The aim is to transcend simple domain categorization by incorporating specific annotations like sub-domains of machine learning, algorithms used, and other relevant information. In this study, we implemented transformer-based natural language processing models to identify and annotate named entities in machine learning articles on arXiv. BERT, in particular, proved to be exceptionally effective, offering high precision in annotation. The methods used and the results obtained highlight the efficiency of these advanced models in enriching academic resources, thus contributing to the creation of an informative and contextual tool for the academic community.