Corpus Construction of Critical Illness Entities and Relationships
摘要
Entity and relational corpus construction is a key part of information extraction and knowledge graph construction. Based on the existing norms of medical entity relationship at home and abroad, we established a disease-centered entity and relationship classification schema according to the characteristics of examination and treatment of cancer-related diseases under the guidance of medical experts. Combined with dictionary, rules, T-Roberta-BiLSTM-CRF entity recognition model and RoBERTa-GSI-PM relation extraction model, we annotate medical texts from multiple sources through multiple rounds of iteration. A Critical Illness entities and relationships Corpus (CIC) was constructed, guided by professional doctors throughout the process, and regular spot checks and consistency checks were taken to ensure the quality of the corpus. Finally, a corpus containing 64,735 entities and 47,222 triplets was constructed, and the consistency of entity and relation annotation reached 0.84 and 0.93, respectively. This corpus provides data basis for medical text information extraction of critical diseases and further research on a series of medical knowledge graph applications.