Topic models are a popular tool for clustering and analyzing textual data. Despite their widespread use in research and application, in-depth analysis methods for topic models are still an important subject of research. Modern methods for interpreting topic models are based on simple visualizations, such as similarity matrices, top-term lists or low-dimensional embeddings, which can only be explained to a limited extent. In this paper we develop a fundamentally different incidence geometric method for analysing and interpreting topic models. For this, we derive ordinal structures from flat topic models, such as non-negative matrix factorization. These enable the analysis of the topic model in a higher (order) dimension and the possibility of extracting conceptual relationships between several topics at once. Due to the use of conceptual scaling, our approach does not introduce any artificial topical relationships, such as artifacts of feature or dimension compression. We introduce and demonstrate the applicability of our approach based on a topic model derived from a corpus of scientific papers taken from 32 top machine learning venues.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Geometric Structure of Topic Models

  • Johannes Hirth,
  • Tom Hanika

摘要

Topic models are a popular tool for clustering and analyzing textual data. Despite their widespread use in research and application, in-depth analysis methods for topic models are still an important subject of research. Modern methods for interpreting topic models are based on simple visualizations, such as similarity matrices, top-term lists or low-dimensional embeddings, which can only be explained to a limited extent. In this paper we develop a fundamentally different incidence geometric method for analysing and interpreting topic models. For this, we derive ordinal structures from flat topic models, such as non-negative matrix factorization. These enable the analysis of the topic model in a higher (order) dimension and the possibility of extracting conceptual relationships between several topics at once. Due to the use of conceptual scaling, our approach does not introduce any artificial topical relationships, such as artifacts of feature or dimension compression. We introduce and demonstrate the applicability of our approach based on a topic model derived from a corpus of scientific papers taken from 32 top machine learning venues.