High-dimensional large-scale mixed-type data imputation under missing at random
摘要
Missingness in mixed-type variables is commonly encountered in a variety of areas. The requirement of complete observations necessities data imputation when a moderate or large proportion of data is missing. However, inappropriate imputation would downgrade the performance of machine learning algorithms, leading to bad predictions and unreliable statistical inference. For high-dimensional large-scale mixed-type missing data, we develop a computationally efficient imputation method, missing value imputation via generalized factor models (MIG), under missing at random. The proposed MIG method allows missing variables to be of different types, including continuous, binary, and count variables, and are scalable to both data size n and variable dimension p while existing imputation methods rely on restrictive assumptions such as the same type of missing variables, the low dimensionality of variables, and a limited sample size. We explicitly show that the imputation error of the proposed MIG method diminishes to zero with the rate Op(max{n−1/2, p−1/2}) as both n and p tend to infinity. Five real datasets demonstrate the superior empirical performance of the proposed MIG method over existing methods that the average normalized absolute imputation error is reduced by 5.3%−34.l%.