AlzViTGNet: A Hybrid Vision Transformer and Graph Convolutional Network Framework for Multiclass Alzheimer’s Disease Classification from MRI Images
摘要
Alzheimer’s disease (AD) is a major neurodegenerative disorder where early and accurate diagnosis is critical. While deep learning has shown promise in MRI-based AD detection, many approaches rely solely on spatial features and overlook anatomical relationships. We propose AlzViTGNet, a hybrid framework combining Vision Transformers (ViT) for global spatial dependencies with Graph Convolutional Networks (GCN) for relational reasoning. ViT-extracted patch embeddings are transformed into graph nodes, enabling structured learning across spatially distributed but semantically linked brain regions. This patch-to-graph fusion enhances the model’s ability to capture subtle structural changes across AD stages (non-demented, very mild, mild, and moderate dementia). Evaluated on a balanced OASIS MRI dataset, AlzViTGNet achieved 98.06% accuracy, outperforming conventional baselines while remaining computationally efficient. These results highlight the value of integrating spatial attention with topological modeling to advance MRI-based AD diagnosis.