Background <p>Alzheimer’s disease (AD) is a progressive neurodegenerative disorder that presents challenges for early detection and intervention. Mild cognitive impairment (MCI), a critical precursor of AD, progresses to dementia in a substantial proportion of individuals annually. Genetic factors, particularly single nucleotide polymorphisms (SNPs), play a key role in the pathogenesis of AD, as identified by genome-wide association studies (GWAS). Therefore, we aimed to develop and evaluate predictive models for classifying patients with MCI into high- and low-risk groups for dementia using SNP chip data and machine learning (ML) algorithms.</p> Methods <p>Using data from the Biobank Innovations for Chronic Cerebrovascular Disease with Alzheimer’s Disease Study, we conducted a GWAS to identify dementia-associated SNPs in a Korean population cohort. The SNPs identified were used to train six ML algorithms—random forest (RF), k-nearest neighbor (KNN), artificial neural network (ANN), support vector machine (SVM), XGBoost, and LightGBM to predict dementia risk. Three predictive models were developed using different SNP subsets: Model 1 (54 SNPs, subjective cognitive decline [SCD] vs. AD + Vascular dementia [VD]), Model 2 (60 SNPs, SCD vs. AD), and Model 3 (76 SNPs, union set of SNPs from the AD vs. SCD and AD + VD vs. SCD). Performance was evaluated primarily using AUC and PR-AUC, which summarize discrimination independent of threshold choice. Thresholds were pre-specified within training folds using Youden’s J (balanced sensitivity/specificity) and F1-max (converter-sensitive) criteria, and then applied unchanged to the temporally separated follow-up cohort.</p> Results <p>In repeated cross-validation, boosting models achieved the strongest performance (e.g., Model 3, XGBoost AUC = 0.881 ± 0.074, PR-AUC = 0.924 ± 0.055). Probabilistic outputs were well-calibrated (Brier scores 0.116–0.183), and calibration plots confirmed good agreement between predicted and observed risks. In a temporally separated follow-up cohort (<i>n</i> = 61, 14 converters), discrimination was modest (AUROC approximately 0.45–0.55), reflecting limited power but showing consistent enrichment of events in predicted high-risk groups. Under F1-max thresholds, sensitivity was high (approximately 0.86–0.93) with NPV approximately 0.80–0.92, whereas specificity was modest (approximately 0.19–0.30) and PPV approximately 0.20–0.27, highlighting the trade-off between capturing converters and limiting false positives.</p> Conclusions <p>Our study highlights the potential of integrating genetic data with ML-based approaches for personalized dementia risk assessment. Although performance was modest in temporal validation, these findings support the feasibility of SNP-based ML stratification in Korean MCI populations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A machine learning framework for classifying dementia risk in mild cognitive impairment: evidence from a Korean genome-wide association study cohort

  • Myeongji Cho,
  • Hyo-Jeong Ban,
  • Hye Ryeong Nam,
  • Chang Hee Chu,
  • Jae Pil Jeon,
  • Sang Cheol Kim

摘要

Background

Alzheimer’s disease (AD) is a progressive neurodegenerative disorder that presents challenges for early detection and intervention. Mild cognitive impairment (MCI), a critical precursor of AD, progresses to dementia in a substantial proportion of individuals annually. Genetic factors, particularly single nucleotide polymorphisms (SNPs), play a key role in the pathogenesis of AD, as identified by genome-wide association studies (GWAS). Therefore, we aimed to develop and evaluate predictive models for classifying patients with MCI into high- and low-risk groups for dementia using SNP chip data and machine learning (ML) algorithms.

Methods

Using data from the Biobank Innovations for Chronic Cerebrovascular Disease with Alzheimer’s Disease Study, we conducted a GWAS to identify dementia-associated SNPs in a Korean population cohort. The SNPs identified were used to train six ML algorithms—random forest (RF), k-nearest neighbor (KNN), artificial neural network (ANN), support vector machine (SVM), XGBoost, and LightGBM to predict dementia risk. Three predictive models were developed using different SNP subsets: Model 1 (54 SNPs, subjective cognitive decline [SCD] vs. AD + Vascular dementia [VD]), Model 2 (60 SNPs, SCD vs. AD), and Model 3 (76 SNPs, union set of SNPs from the AD vs. SCD and AD + VD vs. SCD). Performance was evaluated primarily using AUC and PR-AUC, which summarize discrimination independent of threshold choice. Thresholds were pre-specified within training folds using Youden’s J (balanced sensitivity/specificity) and F1-max (converter-sensitive) criteria, and then applied unchanged to the temporally separated follow-up cohort.

Results

In repeated cross-validation, boosting models achieved the strongest performance (e.g., Model 3, XGBoost AUC = 0.881 ± 0.074, PR-AUC = 0.924 ± 0.055). Probabilistic outputs were well-calibrated (Brier scores 0.116–0.183), and calibration plots confirmed good agreement between predicted and observed risks. In a temporally separated follow-up cohort (n = 61, 14 converters), discrimination was modest (AUROC approximately 0.45–0.55), reflecting limited power but showing consistent enrichment of events in predicted high-risk groups. Under F1-max thresholds, sensitivity was high (approximately 0.86–0.93) with NPV approximately 0.80–0.92, whereas specificity was modest (approximately 0.19–0.30) and PPV approximately 0.20–0.27, highlighting the trade-off between capturing converters and limiting false positives.

Conclusions

Our study highlights the potential of integrating genetic data with ML-based approaches for personalized dementia risk assessment. Although performance was modest in temporal validation, these findings support the feasibility of SNP-based ML stratification in Korean MCI populations.