Alleviating Collapsing Problem in Policy Topic Discovery via Soft Clustering-Based Regulation
摘要
Topic modeling aims to extract the core information implicit in natural language texts. To develop an effective topic model, high-quality input are crucial. However, existing policy texts suffer from the long-tail issue, i.e., a large proportion of words are rarely mentioned in the context, which are called long-tail words. As a result, the topic model trained on these texts tends to recommend frequent items, which leads to high similarity and poor usability of the topics obtained by the model, preventing mainstream topic models from being directly applied to policy texts. To address the above challenge in topic modeling of policy texts, in this paper, we proposes the Embedding Soft Clustering-based Regulation Neural Topic Model for policy topic discovery (ESCRTopic). ESCRTopic first uses Auto-Encoder to learn topic representation self-supervisely, then processing a soft clustering-based embedding regularization via Sinkhorn algorithm based transport optimization to mitigate the long-tailed problem from the perspective of clustering. The experimental results show that our method outperforms other state-of-the-arts topic models in terms of the classical evaluation indexes of topic models, and significantly improves the diversity of topic generation results.