Bengali Document Clustering: A Comparative Study of K-Means, K-Means++, Spectral K-Means
摘要
Text clustering is an active research area. It has various applications in information retrieval, text summarization, etc. Most existing works on document clustering have been carried out in the English domain. A few attempts have been made for document clustering in regional languages like Bengali. In this study, we have implemented some state-of-the-art clustering algorithms for Bengali document clustering. For performance evaluation of the document clustering algorithms, we have created a Bengali dataset consisting of 636 documents which are divided into 32 classes. The experimental result shows that the spectral K-means performs the best among the other clustering algorithms implemented by us.