This paper focuses on similarity measures applicable in agglomerative clustering analysis of datasets with binary variables and their impact on resulting clustering solutions. Specifically, it analyzes 65 measures for binary data. The analysis of the influence of selecting a measure on the outputs of clustering analysis is conducted through a simulation study. The mutual similarity of clustering solutions and the quality of clustering solutions resulting from individual measures are evaluated, using both internal and external evaluation criteria, while two clustering methods are used in the experiment. Finally, two groups of measures leading to almost identical clustering solutions were identified. This way the impact of the measure on clustering results is examined and quantified, which an area that has not been sufficiently explored until now.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Similarity Measures for Datasets with Binary Variables and Their Application in Agglomerative Cluster Analysis

  • Jana Cibulková,
  • Hana Řezanková,
  • Zdeněk Šulc,
  • Jaroslav Horníček

摘要

This paper focuses on similarity measures applicable in agglomerative clustering analysis of datasets with binary variables and their impact on resulting clustering solutions. Specifically, it analyzes 65 measures for binary data. The analysis of the influence of selecting a measure on the outputs of clustering analysis is conducted through a simulation study. The mutual similarity of clustering solutions and the quality of clustering solutions resulting from individual measures are evaluated, using both internal and external evaluation criteria, while two clustering methods are used in the experiment. Finally, two groups of measures leading to almost identical clustering solutions were identified. This way the impact of the measure on clustering results is examined and quantified, which an area that has not been sufficiently explored until now.