错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal Subword De-Identification of Medical Speech for Privacy Protection Leveraging ASR and LLMs

  • Chao-Long Huang,
  • Pratham Nandy,
  • Hong-Jie Dai

摘要

This study presents a cross-lingual de-identification framework for medical speech data that integrates Automatic Speech Recognition (ASR) and Large Languages Models (LLMs), validated on a multilingual dataset (Mandarin Chinese, English, and Taiwanese Hokkien). Through systematic experiments, we evaluated the impact of various data sources and fine-tuning strategies on sensitive information identification. Results show that, after initial fine-tuning on Mandarin, even a small amount of domain-specific multilingual data significantly improves de-identification performance compared to using the base model or a single-language fine-tuned model. The proposed approach, incorporating diverse annotation strategies and robust post-processing procedures such as overlap filtering and fuzzy matching, further enhances accuracy and robustness. Importantly, the framework maintains strong performance under limited annotated data, highlighting its potential for real-world clinical data privacy protection and downstream medical AI applications. Future work may extend this framework to additional languages and more complex medical speech contexts, exploring the feasibility and limitations of speech de-identification in specialized clinical domains.