K-means clustering is a fundamental data mining technique. It heavily relies on parameter optimization (number of clusters, initial centers, and distance measures) for accurate and meaningful results. This study addresses the challenge of simultaneously optimizing these three parameters. It introduces two systems based on the enhanced grey wolf optimization (GWO) algorithm, which handles variable-length individuals. Two individual representations were used: index-based and data-based. Additionally, the good-point set (GPS) population initialization technique was used. The systems were evaluated using four popular clustering measures: the Davies-Bouldin Index, Silhouette Coefficient, V-measure, and Adjusted Rand Index. A proposed hybrid measure (Hybrid Score) was also tested on five well-known datasets. Experiments showed promising results. Both enhanced systems effectively optimized the key k-means parameters. They consistently exhibited lower accuracy variations. The proposed hybrid internal-external measure outperformed traditional metrics, with each metric showing specific strengths for different data types.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing K-Means Clustering Selection Parameters Using Modified Grey Wolf Optimization

  • Athraa Qays Obaid,
  • Maytham Alabbas

摘要

K-means clustering is a fundamental data mining technique. It heavily relies on parameter optimization (number of clusters, initial centers, and distance measures) for accurate and meaningful results. This study addresses the challenge of simultaneously optimizing these three parameters. It introduces two systems based on the enhanced grey wolf optimization (GWO) algorithm, which handles variable-length individuals. Two individual representations were used: index-based and data-based. Additionally, the good-point set (GPS) population initialization technique was used. The systems were evaluated using four popular clustering measures: the Davies-Bouldin Index, Silhouette Coefficient, V-measure, and Adjusted Rand Index. A proposed hybrid measure (Hybrid Score) was also tested on five well-known datasets. Experiments showed promising results. Both enhanced systems effectively optimized the key k-means parameters. They consistently exhibited lower accuracy variations. The proposed hybrid internal-external measure outperformed traditional metrics, with each metric showing specific strengths for different data types.