Conditional information entropy-based feature selection for partially labeled heterogeneous data via matrix operation and prediction label using k-nearest neighbor
摘要
Since labeling data is expensive and time-consuming, partially labeled data is very common. Exploring such data can unlock its potential value and reduce reliance on large amounts of missing labels. k-nearest neighbor rule as an instance-based learning method, is highly effective in predicting missing labels when dealing with partially labeled data, offering significant advantages. This paper uses k-nearest neighbor rule to predict labels for partially labeled heterogeneous data and studies feature selection through conditional information entropy and matrix operations. First, the distance function with respect to each type of feature of partially labeled heterogeneous data is defined, and the tolerance classes are constructed. Then, a new method of predicting the labels for partially labeled heterogeneous data is raised. The core idea of this method is to calculate the distance between the target object and other objects, identify the k-nearest neighbors to the target object, and make predictions based on their labels. By predicting the missing labels, the completeness and usability of the label information can be improved. Next, the tolerance relation matrix, tolerance diagonal matrix and decision relation matrix are defined, and some properties of these matrices are obtained. Moreover, the fact that conditional information entropy for partially labeled heterogeneous data is calculated by matrix operations is proved. Finally, a feature selection algorithm based on the matrix form of conditional information entropy is proposed and compared with six existing algorithms. The experimental results confirm that the algorithm performs excellently and possesses strong robustness and adaptability.