Comparing Alzheimer disease phenotype extraction using rule-based natural language processing, GPT-4, Phi-4, LLaMA, and DeepSeek
摘要
Clinical management of Alzheimer disease (AD) leverages information in unstructured narratives. We compared large language models, GPT-4, Phi-4-14b, DeepSeek-R1-Distill-LLaMA-8b, and LLaMA-3.2-3b, against a previous rule-based pipeline for extracting AD-related phenotypes from 100 real-world clinical notes. Evaluated against clinician annotations, GPT-4 performed best across phenotypes (median precision = 0.96, recall = 1, F1 = 0.98), but rule-based natural language processing was most precise (median = 0.97). Model choice depends on context, considering trade-offs between precision, recall, and scalability.