Explainable AI-Driven Analyses of Pima Indian Cohort Data for Cost Effective and Accurate Prediction of Diabetes Mellitus
摘要
Diabetes Mellitus is a prevalent meta-bolic disorder, and its early prediction is crucial for effective disease management and prevention. The objective of this research is to identify the key risk factors that predominantly contribute to Diabetes Mellitus and utilize them for early disease prediction to facilitate timely medical intervention. In pursuit of clinically meaningful and interpretable outcomes, the study adopts an approach guided by principles of Explainable AI, enabling transparency in the prioritization of risk factors based on their predictive influence. In this aspect the research further explores the interrelationship between key and non-key risk factors, with a focus on uncovering confounding effects that impact predictive accuracy. By determining the critical range of both key and non-key risk factors, the study establishes thresholds that correspond to heightened disease prevalence. The ultimate goal is to construct a prediction model that relies on a minimal set of significant features while providing outputs that are comprehensible to clinicians, enhancing trust and usability through interpretability and informed attribution. The proposed methodology consists of three steps: (1) Identification of key risk factors through a three-stage structured statistical framework involving correlation analysis, significance testing, and feature ranking, ensuring that the selected features are both statistically significant and clinically meaningful; (2) Adjusted Odds Ratio analysis to assess the confounding impact of non-key risk factors on key ones, providing a transparent view of how their interactions influence disease prediction and offering interpretable insights into the structure of risk relationships; and (3) Prediction of Diabetes Mellitus using an advanced Machine Learning model, developed with a focus on minimizing the number of predictive variables while maintaining clinical relevance. Throughout this process, the methodology aligns with the principles of Explainable AI by emphasizing interpretability, traceability of decision logic, and the ability to support human understanding of model behavior. Experimental validation is conducted using the Pima Indian Diabetes dataset, where the proposed methodology successfully identifies Glucose, BMI, Age, Pregnancies, and Diabetes Pedigree Function as the key risk factors with the highest predictive impact. With an impact score of 0.4352, Blood Pressure stands out as the most impactful non-key risk factor compared to other non-key risk factors. The model achieves a classification accuracy of 83.95% ± 1.32%, demonstrating its effectiveness in disease prediction. The proposed statistical learning-based approach to key risk factor identification is new, fast, and accurate compared to the state of the art techniques. By capturing relationships between key and non-key risk factors, it accounts for confounding effects that heighten disease risk. The integration of Explainable AI provides insight into model predictions, enabling clinicians to understand the rationale behind diagnostic outcomes. The proposed Machine Learning model enhances predictive precision. In summary, the proposed research work offers a computationally efficient and clinically relevant solution for risk management not only related to diabetes disease but also other diseases, such as cardiovascular disease, chronic kidney disease, etc.