Unsupervised Feature Selection Using Both Similar and Dissimilar Structures
摘要
Unsupervised feature selection method is widely used to handle the rapid increasing complex and high-dimensional sparse data without labels. Many good methods have been proposed in which the relationships between the similar data points are mainly considered. The graph embedding theory is used which occupies a large proportion. Despite their achievements, the existing methods neglect the information from the most dissimilar data. In this paper, we follow the research line of graph embedding and present a novel method for unsupervised feature selection. Two different viewpoints in the positive and negative are used to keep the data structure after feature selection. Besides a Laplacian matrix by which the most similar data structure is kept, we build an additional Laplacian matrix to keep the least similar data structure. Furthermore, an efficient algorithm is designed by virtue of the existing generalized powered iteration method. Extensive experiments on six benchmark data sets are conducted to verify the state-of-the-art effectiveness and superiority of the proposed method.