<p>Given a data set, sparse subset selection (S3) methods can select a subset from the data set to represent it effectively by skillfully utilizing the theory of sparse representation. S3 methods have become one of the important methods for subset selection. However, a common issue faced by existing S3 algorithms is their inability to handle data sets with missing entries. To address this issue, we propose, for the first time, a novel S3 algorithm for data with missing entries (S3-ME). The S3-ME algorithm integrates the low-rank constraint, sparse self-representation, and subset selection into a unified optimization framework. This allows the algorithm to simultaneously complete missing data entries and perform dissimilarity-based subset selection (DS3) of the data, where the dissimilarity measure used in DS3 can be computed using the obtained sparse self-representation. These three steps are interconnected and alternate with each other until convergence is achieved. To enhance the algorithm’s operational efficiency, this process can be streamlined into two separate stages. The proposed algorithm can be solved within the Alternating Direction Method of Multipliers (ADMM) optimization framework. Experimental results demonstrate the algorithm’s effectiveness in recovering missing entries and performing subset selection on both synthetic and real-world data sets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sparse subset selection algorithm for data with missing entries

  • Xiaobin Zhi,
  • Yongfei Li,
  • Yuqing Qiu,
  • Bingzhe Li

摘要

Given a data set, sparse subset selection (S3) methods can select a subset from the data set to represent it effectively by skillfully utilizing the theory of sparse representation. S3 methods have become one of the important methods for subset selection. However, a common issue faced by existing S3 algorithms is their inability to handle data sets with missing entries. To address this issue, we propose, for the first time, a novel S3 algorithm for data with missing entries (S3-ME). The S3-ME algorithm integrates the low-rank constraint, sparse self-representation, and subset selection into a unified optimization framework. This allows the algorithm to simultaneously complete missing data entries and perform dissimilarity-based subset selection (DS3) of the data, where the dissimilarity measure used in DS3 can be computed using the obtained sparse self-representation. These three steps are interconnected and alternate with each other until convergence is achieved. To enhance the algorithm’s operational efficiency, this process can be streamlined into two separate stages. The proposed algorithm can be solved within the Alternating Direction Method of Multipliers (ADMM) optimization framework. Experimental results demonstrate the algorithm’s effectiveness in recovering missing entries and performing subset selection on both synthetic and real-world data sets.