Variational Loss of Random Sampling for Searching Cluster Number
摘要
Estimating the number of clusters is essential for understanding the complexity and features of data, and performing cluster analysis. Existing integration algorithms for estimating the number of clusters are computationally expensive, while the fast convergent algorithms often lack accuracy. This paper proposes the random sampling likelihood clustering algorithm (RSLC), which uses variational loss to measure the sparsity and estimate the number of clusters, cost only O(NCD) each iteration. RSLC transformed algorithm (RSLCT) is further proposed to improve the accuracy and robustness for the circular data distribution. RSLCT capture the trend of circular data, and generate the substitute points to be clustered. Test results demonstrate that the RSLC algorithm is accurate for Gaussian distribution and RSLCT algorithm is effective for capturing the data with the same trend.