<p>Missingness in mixed-type variables is commonly encountered in a variety of areas. The requirement of complete observations necessities data imputation when a moderate or large proportion of data is missing. However, inappropriate imputation would downgrade the performance of machine learning algorithms, leading to bad predictions and unreliable statistical inference. For high-dimensional large-scale mixed-type missing data, we develop a computationally efficient imputation method, missing value imputation via generalized factor models (MIG), under missing at random. The proposed MIG method allows missing variables to be of different types, including continuous, binary, and count variables, and are scalable to both data size <i>n</i> and variable dimension <i>p</i> while existing imputation methods rely on restrictive assumptions such as the same type of missing variables, the low dimensionality of variables, and a limited sample size. We explicitly show that the imputation error of the proposed MIG method diminishes to zero with the rate <i>O</i><sub><i>p</i></sub>(max{<i>n</i><sup>−1/2</sup>, <i>p</i><sup>−1/2</sup>}) as both <i>n</i> and <i>p</i> tend to infinity. Five real datasets demonstrate the superior empirical performance of the proposed MIG method over existing methods that the average normalized absolute imputation error is reduced by 5.3%−34.l%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-dimensional large-scale mixed-type data imputation under missing at random

  • Wei Liu,
  • Guizhen Li,
  • Ling Zhou,
  • Lan Luo

摘要

Missingness in mixed-type variables is commonly encountered in a variety of areas. The requirement of complete observations necessities data imputation when a moderate or large proportion of data is missing. However, inappropriate imputation would downgrade the performance of machine learning algorithms, leading to bad predictions and unreliable statistical inference. For high-dimensional large-scale mixed-type missing data, we develop a computationally efficient imputation method, missing value imputation via generalized factor models (MIG), under missing at random. The proposed MIG method allows missing variables to be of different types, including continuous, binary, and count variables, and are scalable to both data size n and variable dimension p while existing imputation methods rely on restrictive assumptions such as the same type of missing variables, the low dimensionality of variables, and a limited sample size. We explicitly show that the imputation error of the proposed MIG method diminishes to zero with the rate Op(max{n−1/2, p−1/2}) as both n and p tend to infinity. Five real datasets demonstrate the superior empirical performance of the proposed MIG method over existing methods that the average normalized absolute imputation error is reduced by 5.3%−34.l%.