Evaluation of Named Entity Recognition Software Packages for Data Mining
摘要
Abstract
An important area of data analysis is natural language processing technology, including named-entity recognition. A study was conducted to select the best software packages for this task in Russian news text. A corpus of 70 articles from various resources was compiled and packages such as “Natasha”, “SpaCy”, “Stanza” and “DeepPavlov” were compared. Experiments involved both manual and programmatic extraction of named entities, and metric calculations were performed. The research findings indicated that a combination of “Natasha” and “Stanza” can achieve comprehensive entity extraction in Russian news text.