错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Study of Approaches to Short Text Document Clustering

  • N. Meena,
  • D. S. Jayalakshmi,
  • J. Geetha

摘要

With the increasing volume of information and data, both structured and unstructured, clustering has found applications in various domains such as browsing and search engines. Text document clustering focuses on organizing documents into clusters based on their content similarity. Feature extraction techniques are employed to enhance the results, and the similarity between documents is measured to evaluate the quality of the clusters. This paper presents a literature review that concentrates on the utilization of preprocessing techniques, feature extraction, clustering algorithms, and evaluation measures in document clustering. Through an extensive survey, the paper investigates which algorithms yield superior performance in feature extraction, clustering, and evaluation. According to the survey, nature-inspired algorithms like particle swarm optimization (PSO) offer optimal solutions and are particularly effective for feature extraction. Furthermore, the paper suggests employing a hybrid approach to address the clustering problem, utilizing principal component analysis (PCA) for dimensionality reduction.