Enhancing imputation accuracy of weighted K-nearest neighbor algorithm using non-Gaussian kernel functions
摘要
This study investigates the application of a diverse range of kernel functions for computing weights in the weighted k-nearest neighbor imputation algorithm, aiming to enhance imputation accuracy. While prior research has primarily relied on conventional kernels such as the Gaussian kernel, this study explores the potential advantages of alternative kernel functions, including log-normal, exponential, Weibull, and Rayleigh kernels. Through comprehensive simulations and comparative analyses, the findings reveal that these alternative kernels consistently outperform the traditional Gaussian kernel, with a clear performance pattern emerging. The exponential kernel performs best in low-correlation settings, particularly when smaller neighborhood sizes are used. As correlation strength increases, the Weibull kernel becomes dominant, especially with larger values of k, maintaining its advantage even under moderate correlation. While higher missing data percentages slightly intensify these trends, they do not alter the overall kernel hierarchy. These findings suggest that the exponential kernel is ideal for weakly correlated or uncertain data structures, whereas the Weibull kernel is more suitable for datasets with moderate to high correlations, with both consistently outperforming the conventional Gaussian kernel.