Electronic Medical Record (EMR) text notes are a digital version of a patient’s paper chart. It contains a comprehensive record of a patient’s medical history. EMR text notes are designed to streamline healthcare processes, improve accuracy, and enhance patient care by providing easy access to up-to-date patient information for healthcare providers. The AI-cup 2023 competition for privacy protection and standardization of electronic medical records (EMR) has released a dataset annotated with sensitive health information (SHI) and temporal normalization values. This dataset aims to facilitate the development and evaluation of state-of-the-art natural language processing technologies for the task of privacy protection and standardization of EMR text notes. However, we observed that the annotation distribution for different SHI types is highly unbalanced. We, therefore, proposed a large language model (LLM)-powered data augmented approach to generate synthesized training instances to train an LLM based on the Pythia-410 m model released by EleutherAI. Combined with the pattern-based post-processing method, our team, TEAM_3917, achieved macro-F-scores of 0.8155 and 0.8065 for SHI recognition and temporal information normalization, respectively, which were officially ranked fourth during the competition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Sensitive Health Data Recognition and Normalization Through Large Language Model Driven Data Augmentation

  • Chia-Yi Chao,
  • Cheng-Wei Lin

摘要

Electronic Medical Record (EMR) text notes are a digital version of a patient’s paper chart. It contains a comprehensive record of a patient’s medical history. EMR text notes are designed to streamline healthcare processes, improve accuracy, and enhance patient care by providing easy access to up-to-date patient information for healthcare providers. The AI-cup 2023 competition for privacy protection and standardization of electronic medical records (EMR) has released a dataset annotated with sensitive health information (SHI) and temporal normalization values. This dataset aims to facilitate the development and evaluation of state-of-the-art natural language processing technologies for the task of privacy protection and standardization of EMR text notes. However, we observed that the annotation distribution for different SHI types is highly unbalanced. We, therefore, proposed a large language model (LLM)-powered data augmented approach to generate synthesized training instances to train an LLM based on the Pythia-410 m model released by EleutherAI. Combined with the pattern-based post-processing method, our team, TEAM_3917, achieved macro-F-scores of 0.8155 and 0.8065 for SHI recognition and temporal information normalization, respectively, which were officially ranked fourth during the competition.