The present era belongs to the age of digital devices, where everything is going to be digitized. This makes the massive production of data at faster rates and brings Big Data to light. The arrival of big data has influenced and reformed many domains like Healthcare, Bioinformatics, Agriculture, etc. Healthcare big data has gained the attention of various researchers as it aids in the early diagnosis and prediction of diseases. Generally, real-world healthcare data contains missingness and imbalance distribution of classes which makes classification a challenging problem. Therefore, an efficient method is introduced to deal with the problems of data incompleteness and class imbalance in healthcare data. A class mean and standard deviation-based imputation strategy is proposed to impute the missing values within the region of a particular class. Further, an Improved Elephant Herding Optimization (IEHO) algorithm is utilized to apply the optimal weights for improving the imbalanced learning in healthcare data. A new fitness function is proposed which considers the data missingness and class imbalance along with the feature weighting. The optimized \(k\) -Nearest Neighbor ( \(k\) -NN) learning algorithm is used to perform the classification. The performance of the proposed method is evaluated on four healthcare datasets: Pima diabetes, Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival. A comparative analysis with state-of-the-art methods is performed based on four-performance metrics: accuracy, balanced accuracy, g-mean, and f-score. The proposed method obtained higher classification accuracy for Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival datasets than state-of-the-art methods. Experimental results validate the effectiveness of the proposed method in dealing with the missing values and class imbalance problems in healthcare data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Method to Deal with Missing Values and Class Imbalance in Healthcare Data

  • Harpreet Singh,
  • Birmohan Singh,
  • Manpreet Kaur,
  • Suvita Rani

摘要

The present era belongs to the age of digital devices, where everything is going to be digitized. This makes the massive production of data at faster rates and brings Big Data to light. The arrival of big data has influenced and reformed many domains like Healthcare, Bioinformatics, Agriculture, etc. Healthcare big data has gained the attention of various researchers as it aids in the early diagnosis and prediction of diseases. Generally, real-world healthcare data contains missingness and imbalance distribution of classes which makes classification a challenging problem. Therefore, an efficient method is introduced to deal with the problems of data incompleteness and class imbalance in healthcare data. A class mean and standard deviation-based imputation strategy is proposed to impute the missing values within the region of a particular class. Further, an Improved Elephant Herding Optimization (IEHO) algorithm is utilized to apply the optimal weights for improving the imbalanced learning in healthcare data. A new fitness function is proposed which considers the data missingness and class imbalance along with the feature weighting. The optimized \(k\) -Nearest Neighbor ( \(k\) -NN) learning algorithm is used to perform the classification. The performance of the proposed method is evaluated on four healthcare datasets: Pima diabetes, Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival. A comparative analysis with state-of-the-art methods is performed based on four-performance metrics: accuracy, balanced accuracy, g-mean, and f-score. The proposed method obtained higher classification accuracy for Wisconsin Breast Cancer (Original), Hepatitis, and HCC Survival datasets than state-of-the-art methods. Experimental results validate the effectiveness of the proposed method in dealing with the missing values and class imbalance problems in healthcare data.