Topic Modeling for Short Texts via Adaptive P \(\acute{o}\) lya Urn Dirichlet Multinomial Mixture
摘要
Inferring coherent and diverse latent topics from short texts is crucial in topic modeling. Existing approaches leverage the Generalized P \(\acute{o}\) lya Urn (GPU) model to incorporate external knowledge and improve topic modeling performance. While the GPU scheme successfully promotes similarity among words within the same topic, it has two major limitations. Firstly, it assumes that similar words contribute equally to the same topic, disregarding the distinctiveness of different words. Secondly, it assumes that a specific word should have the same promotion across all topics, overlooking the variations in word importance across different topics. To address these limitations, we propose a novel Adaptive P \(\acute{o}\) lya Urn (APU) scheme, which builds topic-word correlation according to the external and local knowledge, and the Adaptive P \(\acute{o}\) lya Urn Dirichlet Multinomial Mixture (APU-DMM) model that uses the topic-word correlation as an adaptive weight to promote topic inference process. Our extensive experimental study on three benchmark datasets shows the superiority of our model in terms of topic coherence and topic diversity over the eight baseline methods (The code is available at https://github.com/ddwangr/APUDMM ).