The analysis of large corpus of text documents requires automated tools to structure the database. In its many forms, topic modeling is a well-studied concept to automatically find a sparse set of topics that can summarize the entire corpus. Topic modeling assumes that the documents are a simple linear combination of certain topics and that no other information is present. In this work, we assume that information in the form of an undirected graph is available and regularize a statistical topic model with this information. We derive a Gibbs sampling technique that employs a Hamiltonian Monte Carlo computational strategy and demonstrate that the graphical structure can qualitatively improve the topic models derived from the Enron email data set.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph-Directed Topic Models of Text Documents

  • Arjuna Flenner,
  • Cristina Garcia-Cardona

摘要

The analysis of large corpus of text documents requires automated tools to structure the database. In its many forms, topic modeling is a well-studied concept to automatically find a sparse set of topics that can summarize the entire corpus. Topic modeling assumes that the documents are a simple linear combination of certain topics and that no other information is present. In this work, we assume that information in the form of an undirected graph is available and regularize a statistical topic model with this information. We derive a Gibbs sampling technique that employs a Hamiltonian Monte Carlo computational strategy and demonstrate that the graphical structure can qualitatively improve the topic models derived from the Enron email data set.