Evolution of diabetes prediction using the fusion of ANOVA, ADASYN technique and XGBoost based on body composition data
摘要
Diabetes is known as a chronic illness with severe consequences. The rising morbidity rates predict a stunning growth in the global diabetes population, approaching 642 million by 2040, implying that one out of every ten people will be affected. This worrying number highlights the critical need for collaborative efforts from industry and academics to accelerate innovation and foster growth in diabetes risk prediction, eventually saving lives. As the frequency of life-threatening diseases, such as diabetes, rises, Medical Decision Support Systems (MDSS) continue to prove their usefulness in supporting healthcare professionals, particularly physicians, in clinical decision-making procedures. Due to the advancement of technology, machine-learning techniques have made headlines in the early prediction of diabetes. In this paper, we employed machine learning techniques and the Analysis of Variance (ANOVA) method to explore associations between regional body fat distribution and diabetes mellitus in a community adult population, aiming to assess predictive capabilities. We used individual standard classifiers and ensemble learning methods to conduct a retrospective analysis of a portion of data based on body composition. To address the class imbalance problem in the target variable, we also applied three oversampling methods to provide more accurate predictions via learning algorithms. The results demonstrate that XGBoost, based on the Adaptive Synthetic Sampling (ADASYN) method, outperforms the state-of-the-art by achieving an accuracy value of 92.04%. This model exhibits more effectiveness for diabetes prediction compared to other models.