Protein sequence-based classification of Alzheimer’s disease using deep learning and attention mechanism
摘要
Alzheimer’s Disease (AD) is a progressive neurodegenerative disorder with increasing global prevalence. Traditional diagnostic techniques often rely on neuroimaging or cerebrospinal fluid analysis, which can be costly, invasive, or limited in availability. This study introduces AD-ProteinNet, a novel deep learning framework that classifies protein sequences as AD-related or healthy, offering a promising alternative on the basis of protein data. The model integrates Bidirectional Long Short-Term Memory and Bidirectional Gated Recurrent Unit with an attention mechanism to capture complex sequence dependencies and highlight informative residues. Protein sequences sourced from NCBI and UniProt are encoded using diverse amino acid feature descriptors, such as Amino Acid Composition, Dipeptide Composition, Tripeptide Composition, Pseudo-Amino Acid Composition, Composition, Transition, Distribution and Quasi-sequence-order to extract physicochemical and structural attributes. Comparative evaluations against six deep learning models, four machine learning models, and two state-of-the-art approaches demonstrated that AD-ProteinNet achieved up to 97.07% accuracy when composite features were used. It was also evaluated using K-fold cross-validation and cross-dataset validation to demonstrate the model’s generalizability and reliability. Other evaluation metrics included precision, recall, F1-score, and MCC. ROC curve analysis and AUC scores further validated the classification ability of the model, while a statistical Confidence Interval test confirmed the reliability of the model. Furthermore, SHapley Additive exPlanations and Local Interpretable Model-Agnostic Explanations were used to enhance model interpretability by identifying the most influential features. AD-ProteinNet provides a cost-effective, scalable, and biologically meaningful computational approach for early-stage AD detection via protein sequences.