Effective text data management is essential in the age of abundant information. In order to improve text document grouping, this study offers a novel hybrid approach that combines discrete differential evolution (DDE) with genetic algorithm (GA). Cluster centroids are started by GA, and then refined by DDE, achieving a balance between global exploration and local optimization. We investigate the application of TF-IDF and Jaccard similarity measures and contrast our approach with industry standard practices. The higher Fowkes–Mallows index and Silhouette score values in the results demonstrate superior clustering performance. In our data-rich era, processing complicated, unstructured text data requires creative solutions. This paper advances the field of text document clustering by providing those.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Genetic Algorithm and Discrete Differential Evolution for Enhanced Text Document Clustering

  • N. Santhosh Ramchander,
  • Nagaratna P. Hegde

摘要

Effective text data management is essential in the age of abundant information. In order to improve text document grouping, this study offers a novel hybrid approach that combines discrete differential evolution (DDE) with genetic algorithm (GA). Cluster centroids are started by GA, and then refined by DDE, achieving a balance between global exploration and local optimization. We investigate the application of TF-IDF and Jaccard similarity measures and contrast our approach with industry standard practices. The higher Fowkes–Mallows index and Silhouette score values in the results demonstrate superior clustering performance. In our data-rich era, processing complicated, unstructured text data requires creative solutions. This paper advances the field of text document clustering by providing those.