Identify Coherent Topics for Short Text Data by Eliminating Background Words via Topic Attention
摘要
Mining for social media has important significance for public opinion monitoring and guidance. Existing topic models rely on a large amount of complex calculations to eliminate the interference of background words so as to obtain coherent topics. For short text data, complex iterative learning is even more necessary to overcome the problem of data sparsity. How to identify coherence topics with low computational complexity is a very meaningful topic in the context of big data. To address this issue, topic attention that refers to the central tendency of all the comments on a topic in different reports is introduced to filter interference factors of topic words co-occurrence network. Then the community structure of the precise subnetwork is partitioned to identify the coherent topics. Empirical results on two different datasets indicate that topic attention can serve as an effective indicator to improve modularity and enhance topic coherence. Furthermore, the correlation, stability, interpretability, diversity of the identified topics, as well as the complexity and efficiency of the model are analyzed and verified.