错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Creating Metaclusters of Topics in Latent Dirichlet Allocation—Comparison of Bottom-Up and Top-Down Approach

  • Peter Madzík,
  • Lukáš Falát

摘要

Latent Dirichlet Allocation (LDA) is one of the most widespread text mining approaches that enables processing unstructured text data and creating thematic groups from them—i.e., topics. Statistically based methods such as perplexity, topic coherence, and others are supposed to help in determining the optimal number of topics. In practice, however, the resulting number of topics is often too large to make it possible to interpret these topics meaningfully and at the same time comprehensively. In this study, we compare two approaches that can be used for creating thematic groups from the text data. The first approach is bottom-up, which is based on respecting the results of statistical procedures to determine the number of topics, and these topics are subsequently connected to broader statistically or interpretatively related areas—metaclusters. The second approach—top-down—is based on the maximum degree of interpretation with a lower number of metaclusters, which are subsequently decomposed into individual topics. The study offers insights into the positives and negatives of both approaches and, using practical examples, compares the reliability of the results and the degree of their interpretability.