This paper introduces an enhanced version of the K-nearest neighbors (KNNs) algorithm designed to address the difficulties associated with missing data imputation in diverse datasets. The proposed method builds upon the traditional KNN approach that enhances imputation accuracy and efficiency. Missing values are pervasive in real-world datasets and can significantly influence the performance of machine learning models. The presented algorithm employs a refined distance metric calculation, considering not only the feature values but also their relevance and significance in the imputation process. Experimental results on benchmark datasets demonstrate that the enhanced KNN algorithm outperforms traditional imputation methods in terms of imputation accuracy while maintaining efficiency within acceptable limits. The proposed algorithm’s versatility is showcased through experiments on datasets from diverse domains, showcasing its efficacy in handling missing values in various contexts. The presented improvements in imputation accuracy and efficiency make the enhanced KNN algorithm a valuable tool for researchers and practitioners dealing with incomplete datasets in machine learning applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced KNN Imputation for Missing Data

  • E. V. Veena,
  • K. P. Pushpalatha

摘要

This paper introduces an enhanced version of the K-nearest neighbors (KNNs) algorithm designed to address the difficulties associated with missing data imputation in diverse datasets. The proposed method builds upon the traditional KNN approach that enhances imputation accuracy and efficiency. Missing values are pervasive in real-world datasets and can significantly influence the performance of machine learning models. The presented algorithm employs a refined distance metric calculation, considering not only the feature values but also their relevance and significance in the imputation process. Experimental results on benchmark datasets demonstrate that the enhanced KNN algorithm outperforms traditional imputation methods in terms of imputation accuracy while maintaining efficiency within acceptable limits. The proposed algorithm’s versatility is showcased through experiments on datasets from diverse domains, showcasing its efficacy in handling missing values in various contexts. The presented improvements in imputation accuracy and efficiency make the enhanced KNN algorithm a valuable tool for researchers and practitioners dealing with incomplete datasets in machine learning applications.