IMICE: An Improved Missing Data Imputation Using Machine Learning
摘要
When researchers use clinical data for studies, they often face the problem of missing data. Missing data can have a detrimental impact on analytical results and introduce bias into the study. Thus, the handling of missing data is required to obtain results from the forecast. In this paper, we have put forward a method for imputing missing values known as IMICE: Improved Multiple Imputation by Chained Equations which is improved from MICE along with Principle Component Analysis data. We also implemented the existing algorithms called Single Center Imputation from Multiple Chained Equations (SICE) and Multiple Imputation with Chained Equations (MICE) to impute missing values and compared our results. We have compared the performance of our method in Diabetes dataset with MICE and SICE algorithms using this dataset. We simulated missing data rates between 1 and 30%, with the missing data being fully random and the Error metrics are Mean Absolute Error and R squared value to examine the performance of the proposed method. The results of the experiments shows that our method is able to obtain less MAE value and also high R2 value after missing data imputation than the other methods for all the missing data percentages. The findings of this study have important implications for researchers who are dealing with missing data in clinical datasets and may benefit from using our proposed method.