In this paper, we discuss clustering methods adapted to high-dimensional data structured in blocks, where the blocks are obtained by grouping the features collected according to specific themes or subspaces of the initial space of features. The proposed approach is based on the Sparse Subspace K-Means (SSKM) method. SSKM is a partition method in the same way as the K-means algorithm, which assigns a sparse weight system to each cluster in the resulting partition. The method is used in a two-level process. The first level consists of applying SSKM to the blocks separately. This provides relevant features for each block. At level 2, SSKM is again applied to the set of features selected at level 1 to determine the final partition of observations and the relevant features that best explain the clusters. The method is evaluated on simulated and real data. The analysis of the results obtained shows that the use of SSKM with two levels produces significantly better results than SSKM with fewer features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Sparse Subspace K-Means

  • Abdoul Wahab Diallo,
  • Mory Ouattara,
  • François Kaly,
  • Ndèye Niang

摘要

In this paper, we discuss clustering methods adapted to high-dimensional data structured in blocks, where the blocks are obtained by grouping the features collected according to specific themes or subspaces of the initial space of features. The proposed approach is based on the Sparse Subspace K-Means (SSKM) method. SSKM is a partition method in the same way as the K-means algorithm, which assigns a sparse weight system to each cluster in the resulting partition. The method is used in a two-level process. The first level consists of applying SSKM to the blocks separately. This provides relevant features for each block. At level 2, SSKM is again applied to the set of features selected at level 1 to determine the final partition of observations and the relevant features that best explain the clusters. The method is evaluated on simulated and real data. The analysis of the results obtained shows that the use of SSKM with two levels produces significantly better results than SSKM with fewer features.