Unlocking Clinical Narrative Text for Secondary Clinical Usage
摘要
The exponential growth of Electronic Health Records (EHRs) offers a wealth of clinical narratives for improving healthcare, but their unstructured nature and variability make Information Retrieval (IR) challenging. Many existing approaches depend on proprietary datasets, limiting reproducibility. While models like BERT and ClinicalBERT excel in tasks like classification, their potential for EHR-based IR—especially in combining structured and unstructured data—remains underutilized. We introduce ClinicalNarr, a new transformer-based model created to tackle challenges in clinical information retrieval. By using the MIMIC-III dataset, which combines structured ICD-coded data with unstructured clinical narratives, ClinicalNarr improves both retrieval accuracy and efficiency. When tested on the MedNLI benchmark, it achieved an impressive 90.5% accuracy, outperforming BlueBERT. Further correlation validation through expert evaluations confirmed its strong semantic relatedness. ClinicalNarr outperformed BlueBERT in information retrieval tasks, achieving a 33% higher hit rate and consistently outperforming it across various metrics. ClinicalNarr addresses important challenges in clinical information retrieval by offering a practical framework for embedding-based methods. This study serves as a benchmark for future research on clinical IR using the MIMIC-III dataset. It enhances clinical decision support systems and improves how healthcare applications use electronic health records (EHRs) for patient care.