Predicting the Valence Rating of Russian Words Using Various Pre-trained Word Embeddings
摘要
In this work, we conducted a comparative testing of 20 sets of pre-trained vectors to computationally estimate valence ratings of words in the Russian language. The word valence was estimated using neural network predictors. A vector representing a word was fed to the input of a multilayer feed-forward neural network that calculated the valence rating of this word. The currently largest Russian dictionary with valence ratings, KartaSlovSent, was used as a source of word valence ratings for training models. The highest accuracy of valence rating estimation was obtained using a set of fasttext vectors trained on the CommonCrawl corpus that includes 103 billion words. Spearman’s correlation coefficient between human ratings and their machine ratings was 0.859. The high estimation accuracy and the large size of the dictionary allows one to use this set of vectors to extrapolate human valence ratings to the widest range of words in the Russian language. It is also worth mentioning 4 sets of vectors presented on the RusVectores project page and trained using the texts of the Araneum Russicum Maximum and Taiga corpora. Despite a significantly smaller size of the training corpus, using these sets of vectors allows obtaining only slightly lower accuracy. The lowest results were obtained for sets of vectors trained using corpora of news texts.