错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable Dense Embedding for Large-Scale Textual Data via Fast Fuzzy Clustering

  • Olzhas Kozbagarov,
  • Rustam Mussabayev,
  • Alexander Krassovitskiy,
  • Nursultan Kuldeyev

摘要

Efficient and interpretable analysis of large-scale textual data is a major challenge in big data. This paper introduces a novel approach for creating interpretable dense embeddings from extensive text corpora. Our method uses FlexiClust, a fast fuzzy clustering algorithm, combined with word co-occurrence analysis to generate semantically rich text embeddings. This approach balances interpretability, simplicity, and efficiency by merging precise word co-occurrence analysis with advanced fuzzy clustering, forming dense, interpretable vector representations. The proposed technique addresses limitations of traditional sparse vectors and complexities of neural network models, offering improvements in text vectorization. It is particularly beneficial for applications such as news aggregation, content recommendation, semantic search, topic modeling, and text classification in large datasets.