In recent years, artificial intelligence (AI) technology has experienced flourishing development, and large-scale language models (LMs) are increasingly recognized as a pivotal direction for the future of AI in the field of intelligent healthcare. In the realm of clinical medical applications, there is an increasing focus on eliminating patient privacy information from textual electronic health records within medical institutions. To address this, we have developed a two-stage fine-tuning procedure which employs the Pythia model for training purposes, enabling the model to effectively extract and normalize patient-related privacy information. Furthermore, ChatGPT is utilized for data augmentation to generate additional training data, enhancing the overall robustness of the system.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Two-Stage Fine-Tuning Procedure to Improve the Performance of Language Models in Sensitive Health Information Recognition and Normalization Tasks

  • Pin-Sen Chiu,
  • Bo-Wei Hou,
  • Yen-Ting Chen,
  • Shi-Hao Huang

摘要

In recent years, artificial intelligence (AI) technology has experienced flourishing development, and large-scale language models (LMs) are increasingly recognized as a pivotal direction for the future of AI in the field of intelligent healthcare. In the realm of clinical medical applications, there is an increasing focus on eliminating patient privacy information from textual electronic health records within medical institutions. To address this, we have developed a two-stage fine-tuning procedure which employs the Pythia model for training purposes, enabling the model to effectively extract and normalize patient-related privacy information. Furthermore, ChatGPT is utilized for data augmentation to generate additional training data, enhancing the overall robustness of the system.