A single-point and length-representation-based model for nested named entity recognition
摘要
Named entity recognition (NER) is one of the core tasks in natural language processing (NLP), yet accurately extracting nested entities remains a formidable challenge. This study proposes a novel span-based method for nested entity recognition, termed SPLR, which incorporates two knowledge embedding strategies—Prior Knowledge Function (PKF) and Token Length Lexical Corpus (TLLC)—to accurately locate entities of varying lengths. Experimental results show that SPLR-PKF achieves F1 scores of 84.2, 86.5, and 79.6 on ACE2004, ACE2005, and GENIA, with nested levels 1–2 scoring 83.2 and 82.0, respectively. In contrast, SPLR-TLLC attains F1 scores of 84.1, 87.5, and 88.4 on the same datasets and 88.9, 87.5, and 70.6 for nested levels 1–3. Moreover, in NST (i.e., for entities with nested relationships of the same category) NER tasks, the SPLR-TLLC model improves performance by 39.5% compared to previous models.