Comprehensive Analysis of Iris Dataset Using K-Mean and Fuzzy K-Mean Clustering Algorithm
摘要
This paper investigates clustering strategies employed on the Iris dataset, with a specific emphasis on the K-means and Fuzzy K-means algorithms. The procedure starts by employing K-means clustering, which involves normalizing the data and creating distinct groups for different values of k. Visualizations, such as scatter plots and centroid markers, depict the geographical distribution of data points and the central tendencies within clusters. Fuzzy K-means clustering, as a subsequent method, offers a sophisticated methodology that permits data points to be assigned to different clusters with varying levels of membership. Visualizations for Fuzzy K-means, such as scatter plots illustrating fuzzy clusters and medoids, improve the comprehension of clustering results. The elbow approach is utilized to ascertain the ideal number of clusters by graphically illustrating the correlation between the number of clusters and the variance within each cluster. Silhouette ratings provide a quantitative assessment of the quality of clusters. The research demonstrates that K-means consistently produces better silhouette scores than Fuzzy K-means for various cluster counts. This suggests that K-means outperforms Fuzzy K-means in terms of cluster coherence and uniqueness on the standardized Iris dataset. Nevertheless, the selection of algorithms relies on the specific properties of the dataset and the objectives of clustering, highlighting the need of thoroughly considering variables like interpretability and noise tolerance when choosing the most suitable clustering approach.