ESert: An Enhanced Span-Based Model for Measurable Quantitative Information Extraction from Medical Texts
摘要
Measurable quantitative information is one of typical quantitative information in unstructured texts. It consists of entities related to numerics, units and their relationships. How to accurately and efficiently extract measurable quantitative information from unstructured texts remains a challenge. This paper aims to propose an Enhanced span-based joint model named ESert, which uses an n-gram encoder to identify word boundaries with a positive and negative sampling mechanism to extract measurable quantitative information. The experiments evaluate the ESert based on a clinical quantification information dataset containing 1359 Chinese electronic medical records. The results show that our model achieves F1 scores of 97.97% and 97.28% in measurable quantitative information recognition and association, respectively, indicating the effectiveness of the proposed model.