A fundamental difficulty in the field of document clustering is establishing the ideal number of clusters to efficiently categorize texts. In this study, we provide a new model called the nonparametric Bayesian Hierarchical Dirichlet Mixture Model (BHDMM). A nonparametric Bayesian model, unlike typical parametric models, has the unique capacity to construct an endless number of clusters without the requirement to predefine a preset cluster count. To demonstrate the model’s capabilities, we deployed it to a dataset of vehicle customer ratings. We performed a cluster analysis using the nonparametric Bayesian Hierarchical Dirichlet Mixture Model, which revealed interesting patterns and structures in the data. This paper details our technique and findings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Study of the Nonparametric Bayesian Hierarchical Dirichlet Mixture Model for Document Clustering

  • Ratnam Dodda,
  • A. Suresh Babu

摘要

A fundamental difficulty in the field of document clustering is establishing the ideal number of clusters to efficiently categorize texts. In this study, we provide a new model called the nonparametric Bayesian Hierarchical Dirichlet Mixture Model (BHDMM). A nonparametric Bayesian model, unlike typical parametric models, has the unique capacity to construct an endless number of clusters without the requirement to predefine a preset cluster count. To demonstrate the model’s capabilities, we deployed it to a dataset of vehicle customer ratings. We performed a cluster analysis using the nonparametric Bayesian Hierarchical Dirichlet Mixture Model, which revealed interesting patterns and structures in the data. This paper details our technique and findings.