Deidentification and Temporal Normalization of the Electronic Health Record Notes Using Large Language Models: The 2023 SREDH/AI-Cup Competition for Deidentification of Sensitive Health Information
摘要
Electronic Medical Records (EMR) implementation benefits the medical industry with streamlined data analysis, increased patient medication safety, reduced expenses for pathology report storage, and improved medical care efficiency. The EMR text notes hold a patient’s clinical record, comprising notes written by the medical staff and further analyzed by the doctor following diagnosis. However, utilizing the EMR text note in its raw form can expose sensitive personal data belonging to patients and medical personnel. Therefore, safeguarding this private information is of utmost importance. Furthermore, the expression of time information in EMR text notes varies across institutions, which can significantly impact the accuracy and reliability of temporal information analysis. Therefore, normalizing temporal information is also a critical issue. The study presents a competition titled Privacy Protection and Standardization of Electronic Medical Record Competition that addresses recognizing Sensitive health information (SHI) recorded in EMR text notes and normalizing temporal information that poses a risk of identity theft. The competition released a corpus containing synthesized SHIs and normalized temporal information. The highest performance for the SHI recognition (subtask 1) and temporal information normalization (subtask 2) are micro-/macro-F of 0.949/0.912 and 0.844/0.869, respectively. Overall, the average micro/macro score for subtasks 1 and 2 were 0.666/0.496 and 0.6/0.394.