Study of Clustering Algorithms for Mixed-Type Data in Presence of Errors and Correlation
摘要
Clustering mixed-type datasets containing continuous, ordinal, nominal, and binary variables poses significant challenges, especially in the presence of measurement error (ME) and misclassification (Mi). This study examines the impact of ME and Mi on the clustering results of algorithms designed specifically for mixed-type data in the case of data with correlation. The clustering algorithms under evaluation are k-prototypes, Modha-Spangler, KAMILA, HyDaP, and PDQ. By highlighting the influence of data inaccuracies on these methods, our research aims to underscore the importance of choosing robust clustering techniques for analyzing complex, real-world information. The result of this study not only sheds light on the resilience of each algorithm to ME and Mi, but also guides practitioners in selecting the most appropriate clustering methods for their specific data challenges.