MACL: Metric and Attribute Space Co-learning for Qualitative Data Clustering
摘要
Cluster analysis of unlabeled categorical data is crucial in many practical applications, such as medical data analysis, financial risk warnings, recommendation systems, etc. Compared with numerical data in explicit distance space, the adoption of metrics is often critical to the success of cluster analysis on categorical data, where qualitative values do not initially have well-defined similarities. However, categorical data metrics are often defined based on certain prior knowledge with limitations, and a particular metric usually cannot reasonably serve the clustering on different datasets. Furthermore, without a well-established metric space, advanced downstream processing probably cannot function properly to enhance clustering performance. This paper, therefore, first proposes to learn a fusion of metrics that complement each other and then learns to adapt the fusion to clustering tasks for a more appropriate exploration of clusters. Experiments on real public datasets from various domains illustrate the superiority and stability of the proposed method in categorical data clustering.