In this work, we introduce Gradual Information Selection Attention (GISA), a self-attention mechanism tailored for multiple-choice question (MCQ) difficulty estimation. Unlike standard self-attention, GISA iteratively refines token importance over multiple steps, ensuring a progressive, adaptive, and structured focus on critical question components while dynamically filtering irrelevant information. By integrating Entropy-Minimizing Self-Attention (EMSA) to sharpen token selection, Sequential Decision-Based Refinement (SDBR) to stabilize attention updates, and Iterative Contextual Masking (ICM) to suppress misleading distractors, GISA effectively models difficulty estimation in MCQs. Experimental results on Ext-MCQ, TEEMIL-H, TEEMIL-K, and RACE++ demonstrate that GISA, when integrated with Transformer-based models such as BERT, mBERT, and IndicBERT, consistently improves F1-scores by an average of 3% over existing state-of-the-art approaches. Furthermore, an ablation study highlights GISA’s robustness in adapting attention dynamically, even when pretrained weights remain frozen, reinforcing its effectiveness in MCQ difficulty estimation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GISA: Gradual Information Selection Attention for MCQ Difficulty Estimation

  • Manikandan Ravikiran,
  • Tarun Sharma,
  • Arnav Bhavsar,
  • Rohit Saluja

摘要

In this work, we introduce Gradual Information Selection Attention (GISA), a self-attention mechanism tailored for multiple-choice question (MCQ) difficulty estimation. Unlike standard self-attention, GISA iteratively refines token importance over multiple steps, ensuring a progressive, adaptive, and structured focus on critical question components while dynamically filtering irrelevant information. By integrating Entropy-Minimizing Self-Attention (EMSA) to sharpen token selection, Sequential Decision-Based Refinement (SDBR) to stabilize attention updates, and Iterative Contextual Masking (ICM) to suppress misleading distractors, GISA effectively models difficulty estimation in MCQs. Experimental results on Ext-MCQ, TEEMIL-H, TEEMIL-K, and RACE++ demonstrate that GISA, when integrated with Transformer-based models such as BERT, mBERT, and IndicBERT, consistently improves F1-scores by an average of 3% over existing state-of-the-art approaches. Furthermore, an ablation study highlights GISA’s robustness in adapting attention dynamically, even when pretrained weights remain frozen, reinforcing its effectiveness in MCQ difficulty estimation.