Assessing the Text Readability by Use of Language Model Embeddings
摘要
This paper focuses on the development of methods for determining text readability in the Polish language. To achieve this, an extensive user study was prepared, resulting in the creation of a dataset consisting of over 23,000 filled Cloze-tests, along with additional data serving as indicators of text difficulty. This allows for a comprehensive evaluation of readability models conducted, including Bag of Words with Linear Regression, Random Forest, SVR with various extracted text features, three pre-trained Polish Roberta models, and five different word embeddings. All evaluations were correlated with three text difficulty measures computed on our corpus of filled-in cloze tests. Surprisingly, the Fasttext model based on word embeddings achieved the best performance, surpassing the Pisarek reference index by 73.9% and even outperforming the RoBERTa model by 63%.