An Efficient Method to Deal with Missing Values and Class Imbalance in Healthcare Data
摘要
The present era belongs to the age of digital devices, where everything is going to be digitized. This makes the massive production of data at faster rates and brings Big Data to light. The arrival of big data has influenced and reformed many domains like Healthcare, Bioinformatics, Agriculture, etc. Healthcare big data has gained the attention of various researchers as it aids in the early diagnosis and prediction of diseases. Generally, real-world healthcare data contains missingness and imbalance distribution of classes which makes classification a challenging problem. Therefore, an efficient method is introduced to deal with the problems of data incompleteness and class imbalance in healthcare data. A class mean and standard deviation-based imputation strategy is proposed to impute the missing values within the region of a particular class. Further, an Improved Elephant Herding Optimization (IEHO) algorithm is utilized to apply the optimal weights for improving the imbalanced learning in healthcare data. A new fitness function is proposed which considers the data missingness and class imbalance along with the feature weighting. The optimized \(k\) -Nearest Neighbor ( \(k\) -NN) learning algorithm is used to perform the classification. The performance of the proposed method is evaluated on four healthcare datasets: Pima diabetes, Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival. A comparative analysis with state-of-the-art methods is performed based on four-performance metrics: accuracy, balanced accuracy, g-mean, and f-score. The proposed method obtained higher classification accuracy for Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival datasets than state-of-the-art methods. Experimental results validate the effectiveness of the proposed method in dealing with the missing values and class imbalance problems in healthcare data.