<p>Alzheimer’s Disease (AD) represents a growing global health challenge, driven by complex genetic factors and diverse risk contributors. Currently, an estimated 55 million people worldwide are affected by dementia, with AD responsible for 60–70% of these cases. This paper explores the application of advanced machine learning approaches to predict AD risk using Genome-Wide Association Studies data from multiple cohorts, with a particular focus on transfer learning and feature selection techniques. We evaluate the performance of Wide and Deep Neural Networks and Multi-Head Attention in assessing their ability to generalise across datasets. As part of this, we explore knowledge distillation as a strategy to enhance model efficiency through improved generalisation performance in smaller architectures by transferring knowledge from high-capacity models to lightweight ones. Furthermore, the performance of these deep learning approaches is compared with tree-based ensembles, including Random Forest and XGBoost. Our experiments evaluate the generalisability, transferability, and efficiency of these models across different transfer learning scenarios. Findings indicate that aggregating multi-cohort training data significantly enhances predictive performance, highlighting the importance of data diversity in improving AD risk assessment. The proposed knowledge distillation approach enables the transfer of knowledge from a complex teacher model to a simpler student model, significantly improving performance. To enhance interpretability, we apply SHAP (SHapley Additive exPlanations) to the student models, revealing cohort-specific differences in SNP importance and highlighting variants in genes such as ABI3BP and SYN3, both of which are linked to immune and synaptic functions in AD. The integration of SHAP enables transparent interpretation of model decisions and supports the identification of transferable genetic markers, reinforcing the clinical relevance of our framework in AD risk prediction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-cohort genetic risk prediction for Alzheimer’s disease: a transfer learning approach using GWAS and deep learning models

  • Isibor Kennedy Ihianle,
  • Wathsala Samarasekara,
  • Keeley Brookes,
  • Pedro Machado

摘要

Alzheimer’s Disease (AD) represents a growing global health challenge, driven by complex genetic factors and diverse risk contributors. Currently, an estimated 55 million people worldwide are affected by dementia, with AD responsible for 60–70% of these cases. This paper explores the application of advanced machine learning approaches to predict AD risk using Genome-Wide Association Studies data from multiple cohorts, with a particular focus on transfer learning and feature selection techniques. We evaluate the performance of Wide and Deep Neural Networks and Multi-Head Attention in assessing their ability to generalise across datasets. As part of this, we explore knowledge distillation as a strategy to enhance model efficiency through improved generalisation performance in smaller architectures by transferring knowledge from high-capacity models to lightweight ones. Furthermore, the performance of these deep learning approaches is compared with tree-based ensembles, including Random Forest and XGBoost. Our experiments evaluate the generalisability, transferability, and efficiency of these models across different transfer learning scenarios. Findings indicate that aggregating multi-cohort training data significantly enhances predictive performance, highlighting the importance of data diversity in improving AD risk assessment. The proposed knowledge distillation approach enables the transfer of knowledge from a complex teacher model to a simpler student model, significantly improving performance. To enhance interpretability, we apply SHAP (SHapley Additive exPlanations) to the student models, revealing cohort-specific differences in SNP importance and highlighting variants in genes such as ABI3BP and SYN3, both of which are linked to immune and synaptic functions in AD. The integration of SHAP enables transparent interpretation of model decisions and supports the identification of transferable genetic markers, reinforcing the clinical relevance of our framework in AD risk prediction.