Anchor Words Inference for Stochastic Matrix Factorization
摘要
Topic modeling offers a useful way to examine the topic labels of extensive document collections, facilitating the organization and outline of the themes within that collection. Previous researchers have suggested considering the probabilistic model, where each document is the convex combination of topic vectors, and the topic vector is a distribution of words. However, finding an appropriate distribution vector for each topic is not easy for a high-dimensional word co-occurrence space. This work provides an alternative topic vector inference method combined with non-negative matrix factorization for learning high-quality topics. To verify the effectiveness and priority of the proposed method, we experiment with three public benchmark datasets, NIPS, Movies, and NYtimes, and show a competitive performance.