Clustering Ensembles: A Data Perspective Survey
摘要
Clustering ensembles combine multiple base partitions to enhance the robustness, stability, and quality of unsupervised learning. Existing surveys have provided valuable overviews of ensemble generation and consensus functions, but they typically focus on algorithmic taxonomies and pay less attention to how data characteristics shape design choices. At the same time, modern applications increasingly involve large-scale, high-dimensional, multi-view, and dynamic data streams, where naive reuse of classical ensemble techniques may fail. In this survey, we revisit clustering ensembles from a data aware perspective. We organize the literature around a pipeline-based framework with four stages: generation of base clusterings, selection of base clusterings, ensemble representation, and consensus fusion. For each stage, we discuss how key data characteristics, including data type and representation, scale and structure, supervision and external signals, as well as dynamics and resource constraints, affect algorithm design. We further provide representative choices of ensemble components across common data regimes and highlight open challenges for building more data aware clustering ensembles.