A Two-Stage Generative Framework for Sensitive Health Information Extraction and Temporal Normalization in Medical Records
摘要
This study evaluates the performance of automatic speech recognition and sensitive personal information extraction models in medical speech, with a particular focus on improving the accuracy of sensitive health in-formation extraction and timestamp normalization. We adopt a two-stage processing framework in which the Whisper-Large-v3 model is used for speech-to-text conversion, with speech recognition accuracy further improved through noise reduction techniques. Subsequently, we use the Mistral-7B-Instruct pre-trained model, combined with Low-Rank Adaptation fine-tuning, to enhance sensitive information extraction. To improve annotation accuracy, ChatGPT is introduced for non-parametric semantic extraction. Experimental results demonstrate that the generative model outperforms traditional discriminative models in SHI extraction tasks. The results indicate that the proposed method performs well in both sensitive information extraction and timestamp normalization, supporting AI applications in healthcare data privacy protection and management.