<p>Lexical sophistication has garnered attention across diverse research domains in which language production and text complexity are relevant areas of study. Nevertheless, among the myriad existing lexical sophistication measures, the vast majority do not systematically differentiate different senses of polysemous words but rather treat all senses of a polysemous word as equally sophisticated. To address this limitation, the current study introduces a system that automatically assigns the words in a text to CEFR (i.e., the Common European Framework of Reference for Languages) levels based on their senses used in context, using the English Vocabulary Profile as a reference. We further propose a set of fine-grained sense-aware lexical sophistication indices based on the CEFR levels of word senses and evaluate the extent to which these indices can predict holistic scores of second language (L2) English writing quality using 1,236 exam scripts from the CLC-FCE dataset (Yannakoudakis et al., <CitationRef CitationID="CR35">2011</CitationRef>). The results show that these fine-grained sense-aware indices are more strongly correlated with scores than existing lexical sophistication measures, with three significant predictors explaining 11.8% of the variance in holistic scores. A regression model that combines the new indices with existing ones achieves substantially greater predictive power than models built with either set of indices alone. We discuss the potential implications of our findings for future research in L2 lexical sophistication.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing fine-grained sense-aware lexical sophistication indices based on the CEFR levels of word senses

  • Nan Hu,
  • Xiaofei Lu,
  • Renfen Hu

摘要

Lexical sophistication has garnered attention across diverse research domains in which language production and text complexity are relevant areas of study. Nevertheless, among the myriad existing lexical sophistication measures, the vast majority do not systematically differentiate different senses of polysemous words but rather treat all senses of a polysemous word as equally sophisticated. To address this limitation, the current study introduces a system that automatically assigns the words in a text to CEFR (i.e., the Common European Framework of Reference for Languages) levels based on their senses used in context, using the English Vocabulary Profile as a reference. We further propose a set of fine-grained sense-aware lexical sophistication indices based on the CEFR levels of word senses and evaluate the extent to which these indices can predict holistic scores of second language (L2) English writing quality using 1,236 exam scripts from the CLC-FCE dataset (Yannakoudakis et al., 2011). The results show that these fine-grained sense-aware indices are more strongly correlated with scores than existing lexical sophistication measures, with three significant predictors explaining 11.8% of the variance in holistic scores. A regression model that combines the new indices with existing ones achieves substantially greater predictive power than models built with either set of indices alone. We discuss the potential implications of our findings for future research in L2 lexical sophistication.