Uncovering Differential Sensitivity Toward Linguistic Features of Cohesion in Large Language Models
摘要
In recent years, Large Language Models (LLMs) have become widespread in language assessment due to their speed, reliability, and economy. However, questions remain regarding their construct validity. While LLMs have been shown to approach human levels of reliability, the decisions they make are difficult to interpret and it is unknown if their decisions are equally sensitive to the same linguistic features as human decisions. One construct of interest in language assessment is text cohesion, or the extent to which ideas are connected to help readers develop coherent mental representations. This study examines whether LLMs finetuned to quantify text cohesion rely on specific cohesion features that differ from those to which humans likely attend, a phenomenon that we describe as differential sensitivity. We examine differential sensitivity by investigating interactions between the effects of six linguistic features on the cohesion scores of human and LLM raters and the effect of the rater on the cohesion scores. We find that this method is effective for investigating differential sensitivity between raters, and that LLMs exhibit significant differential sensitivity to four out of six linguistic features. The findings have important implications for understanding the differences in language features attended to by humans and LLMs during language assessment.