An enhanced sparse subset selection model based on adaptive projection distance
摘要
Subset selection focuses on identifying representative samples from a large dataset to produce a data subset that can represent the main features of the original data and also reduce the data size in an effective way. Sparse representation theory puts emphasis on enhancing the data sparsity, while subset selection aims to reduce the data size. Sparse subset selection is an important branch of subset selection that incorporates sparse representation theory into the exploration of subset selection. Sparse modeling representative selection model has demonstrated outstanding performance in data classification and video summarization. Nonetheless, it suffers from the issue of poor typicality of the selected representatives. Although some recent works have addressed this problem to some extent by introducing the concept of clustering coding terms, most of the clustering coding terms constructed by these existing methods are implicit, and the model structure is complicated, inefficient, and difficult to be understood. To overcome these difficulties, we propose an adaptive projection distance-based sparse subset selection model. This model constructs a clustering coding item by adaptively calculating the projection distance between data. It can not only eliminate redundant and noisy features in the original data but also ensure that the selected representatives are more typical. To solve the proposed model, an efficient algorithm is developed by virtue of the alternating direction method of multipliers. Extensive experiments with comparative analysis demonstrate that the proposed model outperforms several state-of-the-art models in various tasks including subset selection, cluster analysis, classification, and face image representative selection.