Robust Bayesian Cluster Enumeration for RES Distributions
摘要
Cluster analysisCluster analysis loosely refers to a collection of unsupervisedUnsupervised learning methods that group similar objects within unlabeled datasets into clusters. Clustering methods are highly useful in a variety of applications. For example, in the medical sciences, identifying clusters may allow for a comprehensive characterization of subgroups of individuals. However, in real-world data, the true cluster structure is often obscured by heavy-tailed noise, artifacts, and outliersOutliers. Discovering the true cluster structure becomes even more challenging when the number of clusters is unknown. This chapter addresses the above challenges and revisits some recent developments on robust statistical model-based cluster analysisCluster analysis. In particular, Bayesian robust cluster enumeration criteria based on real elliptically symmetric (RES) distributionsReal elliptically symmetric (RES) distribution are discussed. Such approaches formulate the problem of estimating the number of clusters as a maximization of the posterior probability of multivariate elliptically distributed candidate models. We show, starting from Bayes’ theorem and using asymptotic approximations, how to come up with robust criteria that possess closed-form expressions. Then, robust clustering algorithmsClustering algorithm are discussed that implement the criteria based on arbitrary RES distributed mixture models, and even M-estimatorsM-estimator. Real-data examples highlight the usefulness of the discussed methods. Links to open-source software packages are provided enabling the reader to explore and apply the methods.