Leveraging Weights of Base Clusters and Clusterings in Ensemble to Consensus Clustering
摘要
Consensus clustering, also referred to as clustering ensemble, is gaining increasing attention for enhancing the efficacy of individual clustering techniques. Many ensemble strategies involve constructing a co-association matrix, reflecting the likelihood that pairs of data points should belong to the same or different clusters based on the underlying clustering. Subsequently, various single clustering methods are applied to the co-association matrix to yield the consensus clustering. While the significance of base clusterings and clusters is often considered equally, their actual performance or quality differs. Although some weighted ensemble clustering methods have been proposed in the literature, the integration of both cluster and clustering level weights to achieve robust consensus is largely unexplored. This paper introduces a simple but an innovative approach to similarity-based consensus clustering, presenting a framework that utilises weights from both base clusterings and their constituent clusters to determine the affinity of two data points to a cluster. The technique demonstrates effective performance across diverse datasets, outperforming several established clustering algorithms.