错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Analytical Perspective of Missing Values in Machine Learning

  • Darshanaben Pandya,
  • Abhijeetsinh Jadeja,
  • Sanjay Gour,
  • Saumil B. Trivedi,
  • Hansaben Haribhai Patel,
  • Pradyumansinh Udaysinh Jadeja

摘要

Dealing with absent data holds paramount importance during the preprocessing phase of machine learning. This paper presents an analytical perspective on addressing missing values, considering their impact on model performance and the various techniques employed for mitigation. The nature of absent data, whether occurring entirely randomly, randomly, or systematically, represents a crucial consideration in data analysis and modeling and exerts an impact on the choice of handling methods. The paper discusses common techniques such as deletion and imputation, emphasizing approaches like Basic imputation techniques such as mean, median, and mode filling, along with regression-based imputation and advanced methodologies like K-nearest neighbors imputation, represent a spectrum of approaches for handling missing data and machine learning-based imputation. Feature engineering strategies, including indicator variables and utilizing missing-ness as a feature, are explored. Domain-specific considerations and the iterative nature of imputation are highlighted. Additionally, the paper emphasizes the importance of robust validation using cross-validation techniques to ensure comprehensive assessment of the selected methodologies’ efficacy, offering insights into their overall performance in addressing missing data challenges, comprehensive understanding of handling missing values, aiming to minimize bias, information loss, and enhance the reliability of machine learning models.