A new clustering algorithm is proposed. It deploys the structure of hierarchical clustering by embedding a new similarity measure between observations. The proposed similarity measure relies on a completely non-parametric approach that involves a rank transformation of the distances. The preliminary results on synthetic data confirm the validity of the proposed method in terms of its ability to capture the predefined cluster structures in the data. The advantages of the proposed algorithm are manifold. It offers a clustering approach robust with respect to outliers and specific pattern in the data due to the rank-based approach. A high level of flexibility is involved, since the rank transformation can be adapted to every distance or ordering method. Moreover, a natural way to perform variable selection is included.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Fully Rank-Based Hierarchical Clustering Approach Using Information Imbalance

  • Elena Ballante,
  • Antonietta Mira,
  • Silvia Figini

摘要

A new clustering algorithm is proposed. It deploys the structure of hierarchical clustering by embedding a new similarity measure between observations. The proposed similarity measure relies on a completely non-parametric approach that involves a rank transformation of the distances. The preliminary results on synthetic data confirm the validity of the proposed method in terms of its ability to capture the predefined cluster structures in the data. The advantages of the proposed algorithm are manifold. It offers a clustering approach robust with respect to outliers and specific pattern in the data due to the rank-based approach. A high level of flexibility is involved, since the rank transformation can be adapted to every distance or ordering method. Moreover, a natural way to perform variable selection is included.