错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Algorithms and Metrics to Perform Clustering

  • Manuel Rubiños,
  • Antonio Díaz-Longueira,
  • Míriam Timiraos,
  • Álvaro Michelena,
  • María Teresa García-Ordás,
  • Héctor Alaiz-Moretón

摘要

This study introduces a novel approach in the field of soft computing, focused on determining optimal clustering algorithms and evaluation criteria for datasets. Using six complex datasets and MatLab R2023a software, especially the evalclusters function, various clustering algorithms such as k-means, agglomerative clustering and Gaussian mixture distribution are evaluated, along with evaluation criteria such as Calinski-Harabasz index, Davies-Bouldin index, gap criterion and silhouette metrics. The methodology focuses on selecting the best algorithm and evaluation criteria to perform automatic clustering of a dataset. It is experimented with two- and four-cluster datasets, challenging due to their specific distributions. The procedure includes data loading, elimination of pre-existing clustering assignments, and the use of MatLab’s evalclusters function to determine the optimal number of clusters and sample assignment, considering a range of clusters to be evaluated. The results highlight how different combinations of algorithms and evaluation criteria handle the clustering complex data, with specific examples showing both effective and ineffective clustering. Comparative tables are presented for each dataset, highlighting the combinations of algorithms and criteria that achieve accurate clustering. In conclusion, combinations of algorithms and criteria are identified that work effectively on most datasets, finally recommending the use of k-means with the Calinski-Harabasz index to obtain the optimal number of clusters. This study not only validates previous findings, but also extends the applications of soft computing in the classification of complex data sets.