Sentence Difficulty in Three Languages: Russian Dataset Compared to Italian and English
摘要
The task of predicting text complexity has been extensively studied in various languages. However, there has been relatively less research interest in predicting complexity at the sentence level specifically in the Russian language. In this paper, we conduct experiments using a novel dataset that includes sentence-level annotations for complexity. Our study focuses on examining simple syntactical features and baseline models, such as graph neural networks and pre-trained language models. Furthermore, we compare our findings with existing datasets in Italian and English.