Background <p>Neurodegenerative diseases, including Alzheimer’s disease (AD), Parkinson’s disease (PD), frontotemporal dementia (FTD), and amyotrophic lateral sclerosis (ALS), pose a growing global health burden with limited early diagnostic tools. Radiomics, which extracts high-dimensional quantitative features from medical images [<CitationRef CitationID="CR1">1</CitationRef>, <CitationRef CitationID="CR2">2</CitationRef>], combined with artificial intelligence (AI) methods, has emerged as a promising approach to enhance diagnostic accuracy and prognostic prediction in neuroimaging. However, no prior systematic review has comprehensively evaluated the methodological quality, reproducibility challenges, and diagnostic performance of AI-driven radiomics studies across multiple neurodegenerative diseases using established quality assessment frameworks.</p> Methods <p>This systematic review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [<CitationRef CitationID="CR3">3</CitationRef>]. A comprehensive literature search was performed across PubMed/MEDLINE, Scopus, Web of Science Core Collection, and Embase databases from January 2017 to March 2026. Grey literature sources including conference proceedings from RSNA, ISMRM, and OHBM were systematically searched but excluded from the final synthesis. Two independent reviewers screened titles, abstracts, and full texts with substantial inter-reviewer agreement (Cohen’s kappa = 0.87). Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies version 2 (QUADAS-2) tool and the Prediction model Risk Of Bias ASsessment Tool (PROBAST). Disagreements were resolved through consensus discussion, with a third reviewer consulted when necessary. Data extraction included study design, imaging modality, radiomic feature extraction methodology, AI/ML algorithm, sample size, performance metrics, validation strategy, and external validation status.</p> Results <p>From 1452 screened records, 60 studies met inclusion criteria and were included in the qualitative synthesis. The majority focused on AD and mild cognitive impairment (MCI) (n = 35, 58%), followed by PD and movement disorders (n = 15, 25%), FTD (n = 6, 10%), and other neurodegenerative conditions (n = 4, 7%). Structural MRI was the most commonly used modality (n = 38, 63%), followed by PET (n = 14, 23%) and SPECT (n = 8, 13%). Support vector machines (n = 22), convolutional neural networks (n = 18), and random forests (n = 12) were the most frequently employed AI methods. Reported area under the receiver operating characteristic curve (AUC) values ranged from 0.75 to 0.98 for AD diagnosis and 0.78 to 0.95 for PD classification. However, quality assessment revealed that only 12 studies (20%) performed external validation, and 28 studies (47%) were rated as having high risk of bias, primarily due to small sample sizes, lack of independent test sets, absence of prospective validation, and inadequate reporting of feature extraction parameters. Stratified analysis revealed that studies employing deep learning methods reported significantly higher AUC values (median 0.91) compared to classical machine learning approaches (median 0.85), though deep learning studies also exhibited higher risk of bias due to greater model complexity relative to sample sizes. Meta-analysis was not feasible due to substantial heterogeneity in imaging protocols, feature extraction pipelines, and outcome definitions.</p> Conclusion <p>AI-driven radiomics demonstrates potential for improving neuroimaging-based diagnosis and prognosis of neurodegenerative diseases. However, the field remains substantially limited by methodological heterogeneity, insufficient external validation (only 20% of studies), high risk of bias (47% of studies), and critical reproducibility challenges including scanner variability, feature instability, and data leakage. The predominantly retrospective, single-center nature of existing evidence limits clinical generalizability. Future research should prioritize multi-center prospective validation with pre-registered protocols, standardized radiomics workflows adhering to Image Biomarker Standardisation Initiative (IBSI) guidelines, rigorous assessment of feature reproducibility across scanners and sites, and integration with multiomics data to facilitate responsible clinical translation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial intelligence-driven radiomics in neuroimaging for neurodegenerative disease diagnosis and prognosis: a systematic review

  • Shih-Shuan Fang,
  • Sheng-Han Chen

摘要

Background

Neurodegenerative diseases, including Alzheimer’s disease (AD), Parkinson’s disease (PD), frontotemporal dementia (FTD), and amyotrophic lateral sclerosis (ALS), pose a growing global health burden with limited early diagnostic tools. Radiomics, which extracts high-dimensional quantitative features from medical images [1, 2], combined with artificial intelligence (AI) methods, has emerged as a promising approach to enhance diagnostic accuracy and prognostic prediction in neuroimaging. However, no prior systematic review has comprehensively evaluated the methodological quality, reproducibility challenges, and diagnostic performance of AI-driven radiomics studies across multiple neurodegenerative diseases using established quality assessment frameworks.

Methods

This systematic review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [3]. A comprehensive literature search was performed across PubMed/MEDLINE, Scopus, Web of Science Core Collection, and Embase databases from January 2017 to March 2026. Grey literature sources including conference proceedings from RSNA, ISMRM, and OHBM were systematically searched but excluded from the final synthesis. Two independent reviewers screened titles, abstracts, and full texts with substantial inter-reviewer agreement (Cohen’s kappa = 0.87). Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies version 2 (QUADAS-2) tool and the Prediction model Risk Of Bias ASsessment Tool (PROBAST). Disagreements were resolved through consensus discussion, with a third reviewer consulted when necessary. Data extraction included study design, imaging modality, radiomic feature extraction methodology, AI/ML algorithm, sample size, performance metrics, validation strategy, and external validation status.

Results

From 1452 screened records, 60 studies met inclusion criteria and were included in the qualitative synthesis. The majority focused on AD and mild cognitive impairment (MCI) (n = 35, 58%), followed by PD and movement disorders (n = 15, 25%), FTD (n = 6, 10%), and other neurodegenerative conditions (n = 4, 7%). Structural MRI was the most commonly used modality (n = 38, 63%), followed by PET (n = 14, 23%) and SPECT (n = 8, 13%). Support vector machines (n = 22), convolutional neural networks (n = 18), and random forests (n = 12) were the most frequently employed AI methods. Reported area under the receiver operating characteristic curve (AUC) values ranged from 0.75 to 0.98 for AD diagnosis and 0.78 to 0.95 for PD classification. However, quality assessment revealed that only 12 studies (20%) performed external validation, and 28 studies (47%) were rated as having high risk of bias, primarily due to small sample sizes, lack of independent test sets, absence of prospective validation, and inadequate reporting of feature extraction parameters. Stratified analysis revealed that studies employing deep learning methods reported significantly higher AUC values (median 0.91) compared to classical machine learning approaches (median 0.85), though deep learning studies also exhibited higher risk of bias due to greater model complexity relative to sample sizes. Meta-analysis was not feasible due to substantial heterogeneity in imaging protocols, feature extraction pipelines, and outcome definitions.

Conclusion

AI-driven radiomics demonstrates potential for improving neuroimaging-based diagnosis and prognosis of neurodegenerative diseases. However, the field remains substantially limited by methodological heterogeneity, insufficient external validation (only 20% of studies), high risk of bias (47% of studies), and critical reproducibility challenges including scanner variability, feature instability, and data leakage. The predominantly retrospective, single-center nature of existing evidence limits clinical generalizability. Future research should prioritize multi-center prospective validation with pre-registered protocols, standardized radiomics workflows adhering to Image Biomarker Standardisation Initiative (IBSI) guidelines, rigorous assessment of feature reproducibility across scanners and sites, and integration with multiomics data to facilitate responsible clinical translation.