A multi-modal agent attention model for alzheimer’s disease diagnosis with structural MRI and clinical data
摘要
Early and accurate diagnosis of Alzheimer’s disease (AD) is crucial for timely intervention. However, single-modal structural magnetic resonance imaging (sMRI) has limited discriminative ability, and existing multi-modal methods often rely on simple feature concatenation or computationally expensive self-attention. This paper proposes a Multi-modal Agent Attention (MMA) model for AD diagnosis using sMRI and the Mini-Mental State Examination (MMSE) score. MMA incorporates three key components: (1) an improved 3D agent attention mechanism that reduces attention complexity from O(N2) to O(N); (2) a dynamic gated fusion strategy that adaptively adjusts the contribution of MRI and MMSE features for each subject; and (3) a joint learning framework that combines classification with cross-modal contrastive learning. Under 5-fold cross-validation experiments on ADNI dataset, MMA achieved 97.34 ± 1.03% accuracy and 0.995 ± 0.005 AUC for AD vs. CN, and 92.57 ± 1.97% accuracy and 0.977 ± 0.007 AUC for AD vs. MCI. External validation on AIBL further provided initial evidence of cross-cohort generalization. These results suggest that the proposed MMA framework can effectively integrate sMRI and MMSE information, particularly for more challenging diagnostic settings.