错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating Variable Selection Techniques Under Missing Data: A Simulation Study

  • Catherine Bain,
  • Dingjing Shi

摘要

Variable selection is one of the most pervasive problems researchers face, especially with the increased ease in data collection arising from online data collection strategies. Machine learning methods such as LASSO and elastic net regression have gained traction in the field but are limited in the types of problems for which they are suitable. As such, researchers have pulled more complex techniques, such as the genetic algorithm, from fields like computer science. Although there is strong support in the literature for the use of each of these methods on complete data (McNeish. Multivar Behav Res 50(5):471–484, 2015. https://doi.org/10.1080/00273171.2015.1036965 ; Schroeders et al. PLoS One 11(11):e0167110, 2016. https://doi.org/10.1371/journal.pone.0167110 ), less is known about their relative performance in the presence of missing data. Using a large-scale Monte Carlo simulation, the performance of the LASSO, Elastic Net, and the genetic algorithm is reviewed, for solving variable selection problems in the presence of ignorable missing data. In particular, this chapter incorporates the state-of-the-art missing data handling technique multiple imputation, into the studied tools. All techniques were found to perform at satisfactory levels (as measured by MSE, precision, false positive rate, and computation time) under MCAR and MAR conditions. The genetic algorithm was seen to be most robust to changes in the data.