HmBlogs: A Comprehensive Corpus and Benchmarking Study for Persian Word Embedding and Language Modeling
摘要
This paper introduces the hmBlogs corpus and presents word embedding and language modeling benchmarks for Persian as a low-resource language. The corpus has been prepared based on a collection of nearly 20 million posts of Persian blogs over a period of about 15 years and includes more than 6.8 billion tokens. hmBlogs is prepared independently for the Persian language, and it is available in both raw and preprocessed forms. In this study, the hmBlogs is compared with some of the most important corpora available in Persian, and the results show the superiority of the hmBlogs corpus over the others due to its diversity and quality. To evaluate the corpus, we created a benchmark for both static and contextual embeddings examining how the corpus influences the quality of word embeddings and language models, using analogy tests as a key measure. For evaluating using static word embedding a new semantic analogy dataset for the Persian language is proposed and some word embeddings trained over hmBlogs and other corpora are evaluated using the new analogy dataset as well as the existing analogy datasets. For evaluating the corpus by contextual models, we have pre-trained BERT-large on the corpus and compared its performance in analogy tests against other existing Persian BERT models. For this part, a multiple-choice analogy test dataset is made. The aforementioned evaluations underscore the significance of the corpus and highlight the impact of various factors such as evaluation datasets, model generation techniques, different hyperparameters, and evaluation methods in such studies.