错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Named Entity Recognition Model for Polish Books

  • Krzysztof Sopyla,
  • Paweł Drozda,
  • Krzysztof Ropiak,
  • Urszula Witkowska,
  • Małgorzata Sieniewicz,
  • Sebastian Jankowski

摘要

The main aim of this paper is to conduct research on machine learning methods that can accelerate the process of evaluating book manuscripts for publication by different publishers. The designed methods will be implemented in the system provided by Literacka LTD company. For this purpose, a study was conducted on the possibility of extracting named entities from full-text books within 18 NER categories. A specially prepared dataset was used for the study, which was created by annotators marking NER from selected excerpts of books. For the evaluation, two available models for the Polish language PolishRoBERTa and CNN-NLP were selected, as well as the author’s PoLitBert model was prepared. The achieved results showed that in terms of recall, the author’s model achieved the best results, detecting NER in 9 out of 18 categories at more than 70% accuracy level. On the other hand, in terms of prediction rate, the CNN-NLP model performed best, being able to process the entire book and extract NER in about 20 s. The CNN-NLP model was implemented in the production version of the system because of its swiftness.