<p>Document clustering remains a fundamental task in information retrieval, yet accurately capturing semantic structure in long and context-rich texts poses persistent challenges. In this paper, we propose SMoC-LC (Segment-based Mixture of Clusters with Late Chunking), a novel clustering framework that addresses two key limitations of prior methods: fixed-length segmentation and hard cluster assignments. Our approach introduces Late Chunking to produce flexible, variable-length text segments using long-context embeddings, and employs Gaussian Mixture Models (GMM) to enable soft-probabilistic clustering. We benchmark SMoC-LC and its variants (SBoC, SBoC-LC, SMoC) on seven datasets spanning different domains and structural complexity, including AGNews, 20News-10K, BBCNews, Reuters-21578, and DBpedia (L1-L3). Results show that SMoC-LC consistently improves clustering quality across accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI), with statistically significant gains observed in complex, hierarchical datasets. Our analysis reveals that Late Chunking is especially beneficial for short, structured documents, while soft clustering excels in ambiguous or multi-topic contexts. These findings underscore the need for adaptable clustering strategies aligned with textual granularity and semantic ambiguity.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Segment-Based Bag of Clusters with Mixture Models and Late Chunking for Document Clustering

  • Quoc-Khang Tran,
  • Nguyen-Khang Pham

摘要

Document clustering remains a fundamental task in information retrieval, yet accurately capturing semantic structure in long and context-rich texts poses persistent challenges. In this paper, we propose SMoC-LC (Segment-based Mixture of Clusters with Late Chunking), a novel clustering framework that addresses two key limitations of prior methods: fixed-length segmentation and hard cluster assignments. Our approach introduces Late Chunking to produce flexible, variable-length text segments using long-context embeddings, and employs Gaussian Mixture Models (GMM) to enable soft-probabilistic clustering. We benchmark SMoC-LC and its variants (SBoC, SBoC-LC, SMoC) on seven datasets spanning different domains and structural complexity, including AGNews, 20News-10K, BBCNews, Reuters-21578, and DBpedia (L1-L3). Results show that SMoC-LC consistently improves clustering quality across accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI), with statistically significant gains observed in complex, hierarchical datasets. Our analysis reveals that Late Chunking is especially beneficial for short, structured documents, while soft clustering excels in ambiguous or multi-topic contexts. These findings underscore the need for adaptable clustering strategies aligned with textual granularity and semantic ambiguity.