WMCTCF: A Wavelet and Multi-scale Convolution Based Transformer Cross-Modal Framework for Early Diagnosis of Alzheimer’s Disease
摘要
Cross-modal deep learning algorithms have achieved significant results in the classification of Alzheimer’s disease (AD). However, current fusion methods primarily rely on 3D-based approaches, whose complex algorithms may introduce redundant features and show limitations in extracting discriminative representations as well as effectively integrating heterogeneous data sources. Furthermore, the limited availability of medical data poses additional constraints on fully leveraging the strengths of Transformer models. To overcome these challenges, we propose WMCTCF, a novel Transformer-based cross-modal deep learning framework specifically designed for early AD detection. The proposed method integrates 2D MRI, clinical data, and genetic information, introducing an innovative strategy for image feature extraction by combining wavelet transform with the Swin-Transformer, which effectively captures both global and local features. Additionally, a multi-scale convolution mechanism combined with a Transformer encoder is designed for non-image modalities to better model long-range dependencies in clinical and genetic data. Experimental results demonstrate that WMCTCF achieves an ACC of 0.9859 and an AUC of 0.9947 on the ADNI dataset, marking a significant performance improvement over traditional 3D algorithms. The proposed WMCTCF provides a novel and effective solution for multi-modal AD diagnosis tasks.