Early detection of type 2 diabetes mellitus (T2DM) is crucial for preventing severe complications but remains challenging due to the disease's complex nature. This study aims to improve T2DM prediction by combining the Pima Indians Diabetes Dataset (PIDD) with a larger Microsoft Azure dataset, totaling 15,768 instances, and applying advanced machine learning (ML) techniques. Comprehensive range of ML algorithms was employed, including traditional methods like decision trees, random forests, and support vector machines, as well as advanced models such as XGBoost, CatBoost, LightGBM, and TabNet. Multiple imputation methods (KNN, Hot Deck, MICE) were explored for missing data handling. Performance was assessed using accuracy, precision, F1-score, and AUC-ROC, with SHAP analysis for interpretability. Ensemble methods and advanced models demonstrated superior performance on the combined dataset, with CatBoost and LightGBM achieving up to 94.39% accuracy. The integration of the larger Azure dataset consistently improved model performance across all algorithms. Hot Deck imputation generally yielded the best results on the combined dataset. This comprehensive approach shows significant potential for improving early T2DM detection, offering promising tools for clinical decision support and enabling more timely interventions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Early Detection of Type 2 Diabetes: A Machine Learning Approach with Combined Clinical Datasets

  • Amani Mankouri,
  • Abdeslem Dennai,
  • Farid Kacimi

摘要

Early detection of type 2 diabetes mellitus (T2DM) is crucial for preventing severe complications but remains challenging due to the disease's complex nature. This study aims to improve T2DM prediction by combining the Pima Indians Diabetes Dataset (PIDD) with a larger Microsoft Azure dataset, totaling 15,768 instances, and applying advanced machine learning (ML) techniques. Comprehensive range of ML algorithms was employed, including traditional methods like decision trees, random forests, and support vector machines, as well as advanced models such as XGBoost, CatBoost, LightGBM, and TabNet. Multiple imputation methods (KNN, Hot Deck, MICE) were explored for missing data handling. Performance was assessed using accuracy, precision, F1-score, and AUC-ROC, with SHAP analysis for interpretability. Ensemble methods and advanced models demonstrated superior performance on the combined dataset, with CatBoost and LightGBM achieving up to 94.39% accuracy. The integration of the larger Azure dataset consistently improved model performance across all algorithms. Hot Deck imputation generally yielded the best results on the combined dataset. This comprehensive approach shows significant potential for improving early T2DM detection, offering promising tools for clinical decision support and enabling more timely interventions.