Objective <p>Ordinary imputation methods may not be able to handle heterogeneous co-missing data, such as the lung function measures from the spirometry test in population-based studies. This work aims to review and evaluate various statistical and machine learning imputation methods for estimating the prevalence of impaired lung function, such as chronic obstructive pulmonary disease, using data from public surveys on aging studies.</p> Materials and methods <p>We examined 70 articles and identified different statistical and machine learning methods used in missing data imputation. We selected and applied samples from a pseudo-population dataset and compared their accuracy in estimating the sample lung disease prevalence.</p> Results <p>Unsupervised learning (clustering) methods improve multiple imputations. The k-prototype method outperforms DBSCAN as it can handle categorical data more effectively. Direct imputations based on the predicted values of random forests and artificial neural networks are unsatisfactory.</p> Conclusion <p>When combined with multiple imputations, the k-prototype clustering method appears to be the most suitable one for imputing missing spirometry values. Even if the imputation functions are not the same as those used in simulation, the k-prototype method can improve the estimates of the MI methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of machine learning methods in the imputation of heterogeneous co-missing data

  • Hon Yiu So,
  • Jinhui Ma,
  • Lauren E. Griffith,
  • Narayanaswamy Balakrishnan

摘要

Objective

Ordinary imputation methods may not be able to handle heterogeneous co-missing data, such as the lung function measures from the spirometry test in population-based studies. This work aims to review and evaluate various statistical and machine learning imputation methods for estimating the prevalence of impaired lung function, such as chronic obstructive pulmonary disease, using data from public surveys on aging studies.

Materials and methods

We examined 70 articles and identified different statistical and machine learning methods used in missing data imputation. We selected and applied samples from a pseudo-population dataset and compared their accuracy in estimating the sample lung disease prevalence.

Results

Unsupervised learning (clustering) methods improve multiple imputations. The k-prototype method outperforms DBSCAN as it can handle categorical data more effectively. Direct imputations based on the predicted values of random forests and artificial neural networks are unsatisfactory.

Conclusion

When combined with multiple imputations, the k-prototype clustering method appears to be the most suitable one for imputing missing spirometry values. Even if the imputation functions are not the same as those used in simulation, the k-prototype method can improve the estimates of the MI methods.