Recently, the widespread application of electronic medical records (EMRs) has made protecting patients’ personal privacy information crucial and highly important. However, the sources of these EMRs are different, resulting in various time formats, and this inconsistency makes historical data challenging to read. Therefore, the key to solving the problem lies in two major issues in medical information: de-identification of sensitive health information (SHI) as defined by the Health Insurance Portability and Accountability Act (HIPAA) and normalization of time according to the ISO 8601 standard. To identify the items in case reports, we used the powerful capabilities of large language models (LLMs) as the foundation. We developed a set of algorithms based on In Context Learning (ICL), which ensures that the model inference process is always in the correct order while allowing the features contained in invalid labels to be utilized. The efficacy of our method has been validated through its exceptional performance in a real-world competition, attaining an macro-F-measure of 93.51% on the de-identification task and remarkably achieving over 92% accuracy for time normalization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Privacy Protection and Standardization of Electronic Medical Records Using Large Language Model

  • Chao-Long Huang,
  • Babam Rianto,
  • Jun-Teng Sun,
  • Zheng-Xin Fu,
  • Chung-Hong Lee

摘要

Recently, the widespread application of electronic medical records (EMRs) has made protecting patients’ personal privacy information crucial and highly important. However, the sources of these EMRs are different, resulting in various time formats, and this inconsistency makes historical data challenging to read. Therefore, the key to solving the problem lies in two major issues in medical information: de-identification of sensitive health information (SHI) as defined by the Health Insurance Portability and Accountability Act (HIPAA) and normalization of time according to the ISO 8601 standard. To identify the items in case reports, we used the powerful capabilities of large language models (LLMs) as the foundation. We developed a set of algorithms based on In Context Learning (ICL), which ensures that the model inference process is always in the correct order while allowing the features contained in invalid labels to be utilized. The efficacy of our method has been validated through its exceptional performance in a real-world competition, attaining an macro-F-measure of 93.51% on the de-identification task and remarkably achieving over 92% accuracy for time normalization.