Optimized Fuzzy Stacked Ensemble Model-Based Diabetes Prediction in Pakistan
摘要
Diabetes mellitus, a chronic disorder characterized by elevated blood glucose levels, is the most prevalent disease in Pakistan and a significant risk factor for various other health conditions. Consequently, accurate forecasting of this disease is increasingly crucial. This study proposes an efficient and innovative predictive ensemble model for the accurate prediction of diabetes in Pakistani patients, employing the stacked ensemble approach, a powerful and widely used technique in machine learning. To achieve the objectives of this research, data from 196 respondents were collected from tertiary care hospitals in Punjab, Pakistan, based on four variables: age, blood sugar random (BSR), creatinine, and urea. The proposed stacked prediction model was built by combining optimized fuzzy c-means (FCM-MA) with four supervised models: naive Bayes (NB), support vector machine (SVM) with linear kernel (SVML), SVM with radial basis function kernel (SVMRBF), and feedforward neural network (NN). In the stacked ensemble method, NB and SVM with linear and RBF kernels were used as base learners, while the feedforward NN was utilized as the metalearner for the proposed predictive stacked model (PSM). The performance of the PSM was compared to that of the individual base learners (NB, SVML, and SVMRBF) using four evaluation metrics: F1-score, precision, recall, and overall accuracy. The individual models, SVML, NB, and SVMRBF, achieved consistent overall accuracies. However, the proposed stacked model (PSM) demonstrated superior overall accuracy, along with exceptional performance across all other evaluation metrics. The results indicate that the optimized stacked ensemble method offers rapid and accurate predictions for diabetes patients. By combining the strengths of multiple individual models, the stacked ensemble approach proves to be a robust technique for generating more accurate final predictions. Additionally, the PSM effectively validated the clustering labels obtained through optimized fuzzy c-means by utilizing significant features.