错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Extracting Named Entities from Russian-Language Documents with Varying Degrees of Structural Clarity

  • M. D. Averina,
  • O. A. Levanova

摘要

Abstract—

This study addresses the task of recognizing named entities in Russian texts using the CRF model. We analyze two datasets: well-structured refinancing documents and loosely structured court transcripts. We test the model with various text features and CRF parameters (optimization algorithms). On average, the best F-measure for well-structured documents is 0.99, while for loosely structured ones, it is 0.86.