Multi-Constraints Feature Selection-Based Cross-Pattern Heterogenous Ensemble Learning Model for Diabetic Mellitus Prediction Under Data-Imbalance and Insufficiency
摘要
The world has witnessed high-pace rise in diabetic cases due to the unhealthy lifestyle, metabolic malfunctions etc. The increasing diabetic cases have resulted numerous health complications like retinopathy, heart diseases, etc. The alarming rate of diabetic mellitus has alarmed industry to achieve a reliable computer aided diagnosis solution for the scalable diagnosis demands. Recently, the different bio-physiological information like insulin, glucose, cholesterol, body mass index, age, pregnancy frequency, etc. have been used by the different machine learning algorithms to detect diabetic mellitus. However, none of the state-of-art addressed the at-hand challenges like data-imbalance, skewed learning, local minima and convergence. The deep learning-based solutions too failed to address the challenges like gradient vanishing, gradient explode and accuracy degradation that makes state-of-arts limited towards run-time scalable demands. Ironically, the existing methods directly process aforesaid bio-physiological features for learning and classification without addressing above stated issues, which makes overall solution suspicious. In addition, the existing methods use merely standalone machine learning classifier to learn and predict diabetic mellitus, and hence lack generalizability under large non-linear feature space. To address such challenges and contribute a robust computer-aided diagnosis, this paper proposed multi-constraints feature selection-based cross-pattern heterogenous ensemble learning method for diabetic mellitus prediction. At first, it performs outlier removal followed by k-NN data imputation to alleviate data missing issue. Subsequently, it performs multi-constraints feature selection by using Wilcoxon rank-sum test, univariate logistic regression, cross-correlation analysis, principal component analysis, information gain and Gini-score that cumulatively select a set of most significant features, which was later resampled by using synthetic minority up-sampling technique (SMOTE) method. The resampled features were normalized by using Min–Max normalization, which are followed by cross-pattern learning oriented heterogenous ensemble learning (CP-HEL). Noticeably, the use of data-imputation and SMOTE resampling alleviated the problem of class-imbalance and skewed learning. While, multi-constraints feature selection technique reduced search space so as to alleviate local minima and convergence problem while reducing computational costs. Normalization technique on the other hand thwarted the challenge of over-fitting, while CP-HEL learner with the base-classifiers like support vector machine, decision tree, k-NN, Naïve Bayes, artificial neural network variants, least square-SVM and extreme learning machine improved prediction accuracy. The simulation showed the highest prediction accuracy of 97.16%, F-Measure of 0.95 and AUC of 0.94, which is higher than the other existing methods.