Prompt Engineering and Post-Processing for Sensitive Health Information Recognition
摘要
This study proposes an automated pipeline for de-identifying sensitive health information from doctor–patient voice recordings by integrating automatic speech recognition, large language models, prompt engineering, and post-processing techniques. The proposed system combines OpenAI’s Whisper for transcription, WhisperX for word-level alignment and speaker diarization, and GPT-4-Turbo with structured prompts for SHI extraction. The proposed pipeline supports both English and Chinese and produces structured outputs with precise timestamps. Addressing real-world challenges such as hallucinations, entity misclassification, and formatting errors through iterative prompt refinement and tailored post-processing, our approach significantly improves SHI extraction accuracy. This practical solution advances the application of artificial-intelligence-driven clinical natural language processing for analyzing unstructured multilingual voice data.