Electronic medical records (EMRs) are often populated with private or confidential information related to patients. In the setting of the hospital environment, fragments of information could be collected across various electronic health record systems that could be used to deduce the true identities of patients mentioned in EMR text notes. Therefore, applying a de-identification procedure is a crucial means of safeguarding privacy, especially when handling EMR text notes. By preventing the identification of individuals and reducing the risks associated with personal information, it contributes to regulatory compliance, facilitates research and data sharing, and mitigates the risk of data misuse. In 2023, the Ministry of Education in Taiwan sponsored a large nationwide competition, AI-CUP 2023-privacy Protection and standardization of EMR Challenge to seek automatic de-identification and standardization solutions. As one of the participating teams, we first tried to apply zero-shot, one- shot, and few-shot configurations based on ChatGPT, but the preliminary experimental results showed that ChatGPT cannot follow the prompt to generate the prediction results of its recognized SHIs based on the official output format. Therefore, we decided to primarily utilize the Pythia-160m-deduped pre-trained model to develop our system. Through multiple experiments, the developed system achieved official micro- and macro-F-scores of 0.744356 and 0.5960788 respectively on the PPSEMR test set.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Applying Language Models for Recognizing and Normalizing Sensitive Information from Electronic Health Records Text Notes

  • Sheng-Xuan Huang,
  • Hung-An Cheng,
  • Zheng-Hao Li

摘要

Electronic medical records (EMRs) are often populated with private or confidential information related to patients. In the setting of the hospital environment, fragments of information could be collected across various electronic health record systems that could be used to deduce the true identities of patients mentioned in EMR text notes. Therefore, applying a de-identification procedure is a crucial means of safeguarding privacy, especially when handling EMR text notes. By preventing the identification of individuals and reducing the risks associated with personal information, it contributes to regulatory compliance, facilitates research and data sharing, and mitigates the risk of data misuse. In 2023, the Ministry of Education in Taiwan sponsored a large nationwide competition, AI-CUP 2023-privacy Protection and standardization of EMR Challenge to seek automatic de-identification and standardization solutions. As one of the participating teams, we first tried to apply zero-shot, one- shot, and few-shot configurations based on ChatGPT, but the preliminary experimental results showed that ChatGPT cannot follow the prompt to generate the prediction results of its recognized SHIs based on the official output format. Therefore, we decided to primarily utilize the Pythia-160m-deduped pre-trained model to develop our system. Through multiple experiments, the developed system achieved official micro- and macro-F-scores of 0.744356 and 0.5960788 respectively on the PPSEMR test set.