High Accuracy, Low Applicability: A Systematic Review of Low-Burden AI-Based Risk Prediction for Diabetes and Hypertension in LMICs
摘要
Type 2 diabetes mellitus (T2DM) and hypertension are major global public health challenges. Their rapidly increasing burden necessitates effective risk stratification and early detection systems. Artificial Intelligence (AI) including machine-learning (ML) approaches have demonstrated promising performance for non-invasive cardiometabolic disease risk stratification. However, evidence on AI-based risk stratification systems that integrate anthropometric, dietary, and lifestyle determinants remains limited.
ObjectiveTo synthesize and critically evaluate AI-based risk prediction models for diabetes and hypertension using primarily low-burden anthropometric, lifestyle, and clinical determinants, with a specific focus on their real-world applicability and implementation readiness in LMIC screening systems.
MethodsThe current systematic review used PRISMA 2020 guidelines. Electronic searches were performed in PubMed, Scopus, and IEEE Xplore for studies published between January 2018 and March 2025. The search strategy included keywords-“artificial intelligence”, “machine learning”, “risk assessment”, “type2 diabetes mellitus”, “hypertension”, “anthropometric measurements”, “dietary factors”, and “lifestyle” combined using Boolean operators.
Results142 records were searched using database searches, of which 11 studies were finalized for the review as per the inclusion and exclusion criteria. The sample sizes for the selected studies ranged from approximately 738 to 818,603. The reported AUC values ranged from 0.68 to 1.00; however, as these estimates derive predominantly from internally validated, single-population datasets, their application to heterogeneous LMIC health-system contexts remains undemonstrated.
ConclusionAlthough AI-based risk prediction models demonstrate moderate to high discriminative performance, their deployment readiness remains constrained by three specific structural gaps: the absence of external validation across independent populations, insufficient benchmarking against established non-invasive screening tools, and negligible integration within existing population-level NCD screening workflows. This highlights a critical gap between algorithmic performance and implementation readiness in LMIC contexts, indicating that current models are not yet suitable for large-scale screening without further validation and system-level integration.