Named entity recognition (NER) is a particularly challenging task, especially for historical documents that lack extensive annotated datasets [6, 14]. The titles of ukiyo-e, a genre of Japanese artworks, contain a significant number of entities and are composed of short texts rich in historical information. The complexity and brevity of these titles pose considerable challenges for analysis. This paper presents the construction of an ukiyo-e NER dataset and introduces a BERT-based NER methodology that achieves notable success in entity recognition within ukiyo-e titles. The study underscores the effectiveness of BERT and its derivative models on the ukiyo-e NER dataset, proposing a viable NER solution for historical documents. The proposed approach demonstrates the capability to perform NER tasks on ukiyo-e titles using pre-trained models on a relatively small annotated dataset, achieving an accuracy exceeding 80%. Furthermore, this paper examines the distinct characteristics of BERT and its derivative models, optimizing their application to the ukiyo-e NER dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A BERT-Based Method of Named Entity Recognition for Ukiyo-e Titles

  • Bohao Wu,
  • Akira Maeda

摘要

Named entity recognition (NER) is a particularly challenging task, especially for historical documents that lack extensive annotated datasets [6, 14]. The titles of ukiyo-e, a genre of Japanese artworks, contain a significant number of entities and are composed of short texts rich in historical information. The complexity and brevity of these titles pose considerable challenges for analysis. This paper presents the construction of an ukiyo-e NER dataset and introduces a BERT-based NER methodology that achieves notable success in entity recognition within ukiyo-e titles. The study underscores the effectiveness of BERT and its derivative models on the ukiyo-e NER dataset, proposing a viable NER solution for historical documents. The proposed approach demonstrates the capability to perform NER tasks on ukiyo-e titles using pre-trained models on a relatively small annotated dataset, achieving an accuracy exceeding 80%. Furthermore, this paper examines the distinct characteristics of BERT and its derivative models, optimizing their application to the ukiyo-e NER dataset.