STSACM: A Method for Segmenting and Classifying Subjects in Chinese Electronic Medical Records
摘要
Subject segmentation of Chinese electronic medical records (CEMR) is an important way to improve the accuracy of information extraction. However, the uneven distribution of subjects in CEMR and the mutual influence between different subjects have led to high difficulty and low accuracy in subject segmentation. In order to improve the accuracy of CEMR subject segmentation, this paper designs a sentence based CEMR subject segmentation model (STSACM). The model On the one hand provides support for extracting deep semantic features of sentences by obtaining contextual information, key information, and vocabulary combination information that express local features within the sentence; On the other hand, by utilizing gate control mechanisms to integrate sentence level contextual information, a sentence level context perception module is established to achieve dynamic detection of sentence level semantic environments and improve the accuracy of topic identification; Finally, using a joint learning strategy and defining a joint loss function that includes focus loss and binary cross entropy loss, the model enhances the coherence between sentence subjects while improving the accuracy of topic segmentation. The model STSACM is able to focus more on samples that are small in quantity and difficult to distinguish during subject classification, thereby improving its ability to handle complex subject classification. The experimental results on the CEMR datasets show that compared with the classic model WIKI-727K, STSACM has improved accuracy, precision, and F1 value by 1.64%, 1.22%, and 2.91%, respectively, indicating its effectiveness in CEMR subject segmentation.