Nested named entity recognition in traditional Chinese medicine electronic medical records via dual-granularity feature augmentation and span classification
摘要
Named Entity Recognition (NER) plays a crucial role in extracting important information such as treatment methods, symptoms, and herbal prescriptions from Traditional Chinese Medicine (TCM) electronic medical records. However, existing NER methods often struggle with the complexity and variability of TCM language, especially when dealing with overlapping or nested entities. To address these issues, we propose DG-SpanTCM, a novel framework that enhances character-level text understanding using a pre-trained language model and improves entity recognition through lexical-semantic features and robust training strategies. Our method also incorporates techniques to handle label imbalance and better identify complex entity structures. Experiments on a real-world TCM dataset show that DG-SpanTCM achieves superior performance, improving the F1-score over strong baseline models. These findings highlight the potential of DG-SpanTCM in advancing automated information extraction for TCM texts.