<p>Assessing the quality of clustering results is an essential part of cluster analysis. Since traditional unsupervised learning methods typically lack clear ground-truth labels for comparison, evaluating their performance can be challenging. Researchers have created various internal cluster validity (CVI) indices to evaluate clustering using predicted labels and data. Constructing an effective CVI without employing ground-truth values is as challenging as developing clustering techniques. Multiple CVIs are essential because no single CVI is suitable for all datasets and there is no definitive method for selecting the right CVI when the actual labels are unavailable. This paper presents a new internal CVI, the Statistical Test-Based Separability Index (STSM). It uses the Anderson–Darling distance to assess how well the data are separated. We compared STSM with 12 existing CVIs, ranging from the initial Dunn (J 1974) to the most recent CVDD (2019), and an external CVI that serves as the ground truth by using the results of 9 different clustering algorithms on 16 real datasets. The findings demonstrate that STSM outperforms existing CVIs in terms of effectiveness and competitiveness. Additionally, we compiled the basic procedure for assessing CVI and developed a rank difference metric for comparing CVI outcomes.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Statistical Test Based Separability Measure for Internal Cluster Validation

  • K. Kalpanarani,
  • G. Hannah Grace

摘要

Assessing the quality of clustering results is an essential part of cluster analysis. Since traditional unsupervised learning methods typically lack clear ground-truth labels for comparison, evaluating their performance can be challenging. Researchers have created various internal cluster validity (CVI) indices to evaluate clustering using predicted labels and data. Constructing an effective CVI without employing ground-truth values is as challenging as developing clustering techniques. Multiple CVIs are essential because no single CVI is suitable for all datasets and there is no definitive method for selecting the right CVI when the actual labels are unavailable. This paper presents a new internal CVI, the Statistical Test-Based Separability Index (STSM). It uses the Anderson–Darling distance to assess how well the data are separated. We compared STSM with 12 existing CVIs, ranging from the initial Dunn (J 1974) to the most recent CVDD (2019), and an external CVI that serves as the ground truth by using the results of 9 different clustering algorithms on 16 real datasets. The findings demonstrate that STSM outperforms existing CVIs in terms of effectiveness and competitiveness. Additionally, we compiled the basic procedure for assessing CVI and developed a rank difference metric for comparing CVI outcomes.