Fair positive unlabeled learning for predicting undiagnosed Alzheimer’s disease in diverse electronic health records
摘要
Alzheimer’s Disease (AD), the most common neurodegenerative disease, is underdiagnosed and more prominent in underrepresented groups. We performed semi-supervised positive unlabeled learning (SSPUL) coupled with racial bias mitigation for equitable prediction of undiagnosed AD from diverse populations at UCLA Health using electronic health records. SSPUL achieved superior sensitivity (0.77–0.81) and area under the precision recall curve (AUCPR) (0.81–0.87) across non-Hispanic white, non-Hispanic African American, Hispanic Latino, and East Asian groups compared to supervised baseline models (sensitivity: 0.39–0.53; AUCPR: 0.3–0.7). SSPUL also exhibited superior fairness as evidenced by the lowest cumulative parity loss. We identified top shared and distinct features among labeled and unlabeled AD patients, including those that are neurological (e.g., memory loss) and non-neurological (e.g., decubitus ulcer). We validated our results using polygenic risk scores, which were higher in labeled and predicted positives than in predicted negatives among non-Hispanic white, Hispanic Latino, and East Asian groups (p < 0.001).