Prompt-optimized self-supervised double-tower contextualized topic model
摘要
Microblog has a large amount of data and is updated quickly with current events. Unsupervised machine learning models are used for the topic clustering. Complex unsupervised microblog data lacks effective training. The number of fixed topics and their content cannot be set and clustered. Semantic dependencies between words are ignored in topic representations. Microblog data is redundant. Chinese words are less clearly delineated and more semantically difficult. Valid features cannot be extracted. To solve the above problems, a Prompt-optimized self-supervised Double-Tower Contextualized Topic Model(PDTCTM) is proposed. A feature extraction model based on a double-tower structure is used to obtain word embeddings. The double-tower structure is a dynamic integrating of the Contextual pre-training model SBERT and the Chinese pre-training model Chinese-RoBERTa. Keywords are introduced into the pre-training models SBERT and Chinese-RoBERTa as prompt information. The optimized double-tower model performs feature vectorization of words and phrases. The CTM topic model is used to sample and vocabulary is generated according to the topic. Topic modeling is completed. The experimental results show that the PDTCTM model improves the PMI and cos