Weakly supervised multi-label feature selection based on shared subspace
摘要
Multi-label feature selection (MLFS) improves classification accuracy and alleviates the curse of dimensionality by retaining relevant features and eliminating redundant and irrelevant features. However, the incompleteness of the label space not only makes the models difficult to learn the latent structure of the label space, but also leads to the unreliability of correlation between the labels and features. Therefore, it is essential and difficult how to discover the credible correlation information between the features and labels, so as to enable the effective selection of the key feature subset by the algorithms learning from multi-label data with missing labels. To more effectively excavate the implicit shared information within the feature matrix and the label matrix, we propose a novel MLFS method named WSMF which combines the feature matrix and the label matrix together to identify vital feature subsets in the absence of a large portion of labeled data. First, we utilize joint matrix factorization to uncover a low-dimensional shared mode between the feature matrix and the label matrix, thereby diminishing the effect of incomplete label information. Second, we employ non-negative matrix factorization (NMF) to enhance the interpretability of the succeeding feature selection process. Additionally, we employ the structural consistency assumption to retrieve the absent labels and incorporate