Document clustering plays a crucial role in various information retrieval tasks. Existing approaches often struggle with capturing the semantic relationships between documents, especially when dealing with long and complex texts. To address this issue, we propose SBoC, a novel Segment-based Bag-of-Clusters approach. SBoC first divides documents into segments, capturing local semantic information. It then applies clustering algorithms to these segments, forming clusters that represent distinct semantic concepts. Finally, a Bag-of-Clusters representation is constructed for each document, encoding its semantic content based on the assigned segment clusters. SBoC shows promising results, particularly in terms of capturing semantic relationships in document clustering. While not surpassing all existing methods, SBoC demonstrates competitive performance on benchmark datasets, particularly when handling long and complex texts. This approach provides a potential solution for enhancing document clustering for various information retrieval tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SBoC: A Segment-Based Bag of Clusters Approach for Document Clustering

  • Quoc-Khang Tran,
  • Nguyen-Khang Pham

摘要

Document clustering plays a crucial role in various information retrieval tasks. Existing approaches often struggle with capturing the semantic relationships between documents, especially when dealing with long and complex texts. To address this issue, we propose SBoC, a novel Segment-based Bag-of-Clusters approach. SBoC first divides documents into segments, capturing local semantic information. It then applies clustering algorithms to these segments, forming clusters that represent distinct semantic concepts. Finally, a Bag-of-Clusters representation is constructed for each document, encoding its semantic content based on the assigned segment clusters. SBoC shows promising results, particularly in terms of capturing semantic relationships in document clustering. While not surpassing all existing methods, SBoC demonstrates competitive performance on benchmark datasets, particularly when handling long and complex texts. This approach provides a potential solution for enhancing document clustering for various information retrieval tasks.