Automatic Readability Assessment Based on Phraseological Complexity
摘要
Lexical measures are important grading features in readability assessment studies. However, these measures are based on single words with few measures at the level of word combinations. This paper constructs a phraseological complexity feature system from a phraseological dimension, which is used to construct machine learning models in Chinese text complexity automatic grading task experiments. Experiments using five models compare the prediction of traditional lexical complexity and phraseological complexity features on Chinese text grading. Results of all the experiments show that the phraseological dimension features are more predictive than the lexical features, proving the important role of the phraseological dimension features in automatic readability assessment.