This research explores the utilization of artificial intelligence (AI) language generation models for the de-identification of medical case narratives, effectively anonymizing patient identifiers to prevent unauthorized disclosure of sensitive health information. This ensures that patient data remains unrecognizable even when accessed, offering two significant advantages: compliance with legal standards forbidding the indiscriminate release of medical records, and bolstered trust in healthcare providers which encourages patients to share necessary details for therapeutic and investigative purposes. Historically, de-identification processes necessitated the manual examination of records by specialists to identify and sanitize personal identifiers, it is a time-consuming task with potential for error. The digitalization of medical records now offers avenues for automated de-identification, simultaneously enhancing data accuracy and facilitating the ethical utilization of extensive datasets in healthcare research, public health strategy, and policymaking. The emergence of sophisticated generative AI technologies has augmented the capacity for comprehensive de-identification. Platforms such as ChatGPT and Bing have evolved to address complex human inquiries. Yet, employing these tools without proper safeguards potentially compromises de-identification objectives by risking data exposure. To advance the integrity of privacy measures and maintain data governance, this study fine-tunes a pre-established vast language AI model, Pythia, for localized training with state-of-the-art technology with rigorous privacy assurance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utilizing Large Language Models for Privacy Protection and Advancing Medical Digitization

  • Zhu-Jian Ru,
  • Omkar Panchal,
  • Ching-Tai Chen,
  • Jitendra Jonnagaddala,
  • Hong-Jie Dai,
  • Sheng-Chun Hsueh

摘要

This research explores the utilization of artificial intelligence (AI) language generation models for the de-identification of medical case narratives, effectively anonymizing patient identifiers to prevent unauthorized disclosure of sensitive health information. This ensures that patient data remains unrecognizable even when accessed, offering two significant advantages: compliance with legal standards forbidding the indiscriminate release of medical records, and bolstered trust in healthcare providers which encourages patients to share necessary details for therapeutic and investigative purposes. Historically, de-identification processes necessitated the manual examination of records by specialists to identify and sanitize personal identifiers, it is a time-consuming task with potential for error. The digitalization of medical records now offers avenues for automated de-identification, simultaneously enhancing data accuracy and facilitating the ethical utilization of extensive datasets in healthcare research, public health strategy, and policymaking. The emergence of sophisticated generative AI technologies has augmented the capacity for comprehensive de-identification. Platforms such as ChatGPT and Bing have evolved to address complex human inquiries. Yet, employing these tools without proper safeguards potentially compromises de-identification objectives by risking data exposure. To advance the integrity of privacy measures and maintain data governance, this study fine-tunes a pre-established vast language AI model, Pythia, for localized training with state-of-the-art technology with rigorous privacy assurance.