错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bengali Document Clustering: A Comparative Study of K-Means, K-Means++, Spectral K-Means

  • Amartya Roy,
  • Kamal Sarkar,
  • Chintan Mandal

摘要

Text clustering is an active research area. It has various applications in information retrieval, text summarization, etc. Most existing works on document clustering have been carried out in the English domain. A few attempts have been made for document clustering in regional languages like Bengali. In this study, we have implemented some state-of-the-art clustering algorithms for Bengali document clustering. For performance evaluation of the document clustering algorithms, we have created a Bengali dataset consisting of 636 documents which are divided into 32 classes. The experimental result shows that the spectral K-means performs the best among the other clustering algorithms implemented by us.