Topic Label Generation in the Popular Science Corpus
摘要
This chapter presents the results of experiments on topic label generation from web data and distributional semantic models. The procedure in question is required for topic label assignment in Russian popular science corpora. Topic modeling is performed by means of a series of algorithms, including Nonnegative matrix factorization, Latent Dirichlet Allocation, and Biterm topic modeling. Our approach allows for reducing the shortcomings of conventional topic label assignment by choosing the first topical term as a topic label. We introduce an improved version of topic label generation as an ensemble of heterogeneous methods. Candidate labels are evaluated in the course of human assessments. The results of our research allow us to verify the structure of scientific media sites and thus to improve their quality.