Accurately estimating item difficulty is critical for building adaptive and fair educational assessments. However, most automated approaches treat difficulty as a static property, ignoring how subtle linguistic variations—such as stem rewording or distractor phrasing—can shift perceived complexity. We introduce ScoreCLIQ, a modular framework for dynamic item difficulty estimation that integrates model feedback with paraphrastic refinement. ScoreCLIQ first trains a BERT-based estimator on labeled difficulty scores. A Large Language Model (LLM) is then optimized via reinforcement learning to paraphrase items with high prediction error, using the frozen estimator’s output as a reward signal. Finally, the estimator is refined using a regularized objective that promotes prediction consistency between original and paraphrased items. This feedback-driven loop requires no additional human annotations and enables ScoreCLIQ to adapt to surface-level variation in item phrasing. On the BEA 2024 shared task dataset, ScoreCLIQ achieves a 12.46% reduction in RMSE over the previous state-of-the-art. Our results demonstrate that linguistic sensitivity in difficulty estimation can be learned through model-guided self-supervision.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ScoreCLIQ: A Dynamic LLM-Based Framework for Item Difficulty Estimation

  • Soujatya Sarkar,
  • Manikandan Ravikiran,
  • Rohit Saluja

摘要

Accurately estimating item difficulty is critical for building adaptive and fair educational assessments. However, most automated approaches treat difficulty as a static property, ignoring how subtle linguistic variations—such as stem rewording or distractor phrasing—can shift perceived complexity. We introduce ScoreCLIQ, a modular framework for dynamic item difficulty estimation that integrates model feedback with paraphrastic refinement. ScoreCLIQ first trains a BERT-based estimator on labeled difficulty scores. A Large Language Model (LLM) is then optimized via reinforcement learning to paraphrase items with high prediction error, using the frozen estimator’s output as a reward signal. Finally, the estimator is refined using a regularized objective that promotes prediction consistency between original and paraphrased items. This feedback-driven loop requires no additional human annotations and enables ScoreCLIQ to adapt to surface-level variation in item phrasing. On the BEA 2024 shared task dataset, ScoreCLIQ achieves a 12.46% reduction in RMSE over the previous state-of-the-art. Our results demonstrate that linguistic sensitivity in difficulty estimation can be learned through model-guided self-supervision.