Extending the Boosted-Oriented Probabilistic Clustering to the Unit Hypersphere: A Textual Data Perspective
摘要
Clustering techniques play a crucial role in uncovering patterns and structures within complex datasets. In this study, we extend the boosted-oriented probabilistic clustering algorithm, originally proposed for time series data, to address the unique challenges posed by directional data. This approach is applied to a textual dataset to categorize documents based on their word content. The results highlight the effectiveness of this technique in organizing high-dimensional, sparse data, offering a novel perspective for clustering in challenging domains.