Background <p>Adapting instruction to learners’ moment-to-moment cognitive load and achievement emotions remains challenging in deployed Intelligent Tutoring Systems (ITS). We present a reliability-aware, real-time multimodal framework that senses facial, vocal, physiological, and interaction signals, estimates joint affect–cognition states, and delivers state-contingent interventions under tight latency constraints.</p> Objective <p>To (i) design an end-to-end platform-agnostic pipeline that integrates Cognitive Load Theory and Control–Value accounts of achievement emotions within an Integrated Cognitive–Affective Learning Model; (ii) validate technical performance and educational impact in laboratory and classroom settings; and (iii) report effect sizes and uncertainty per reviewer guidance.</p> Methods <p>A randomized, mixed-methods evaluation contrasted four conditions (Full adaptive, Cognitive-only, Affective-only, Control). The fusion layer aligns and normalizes modalities (DTW • PCA) with reliability weighting. A temporal Transformer and a GRU estimate emotions and triarchic load. A constrained contextual bandit selects content, interface, metacognitive, or motivational actions. Participants (<i>N</i> = 271) completed math and science units over nine sessions spanning 3 weeks. Primary outcomes included pre- and post-learning assessments, as well as transfer. Process outcomes were measured in terms of time spent in engagement versus confusion, frustration, and boredom. UX and latency were also logged. Mixed-effects models yielded estimates with 95% CIs; sensitivity checks addressed sensor quality and site effects.</p> Results <p>The system met real-time targets (end-to-end latency &lt; 200 ms; reaction time 520–820 ms) while maintaining an affect recognition accuracy of 82–88% (weighted F1 ≈ score 0.85). Adaptive instruction improved post-test performance compared to a strong static ITS baseline, with the most significant gains for mid-ability learners (approximately + 15% points, Cohen’s d approximately 0.8). Process measures showed higher time in Engagement (≈ 42% vs. 29%) and productive Confusion (≈ 18% vs. 12%), and lower Frustration/Boredom, with strong convergence between detected negative affect and self-reported frustration (<i>r</i> ≈ 0.7). Mechanism-level analyses indicated state-contingent efficacy: content adaptations during Confusion, interface adjustments during Frustration, metacognitive prompts during Engagement, and motivational supports during Boredom.</p> Conclusions <p>A theoretically grounded, reliability-aware multimodal controller can translate fine-grained state estimation into timely, pedagogically meaningful adaptations that scale across laboratory and classroom contexts. Benefits concentrate where scaffolding has the highest leverage (mid-ability band) and are mediated by improved regulation of cognitive load and affect. We report limitations (including more challenging anxiety detection, classroom noise, and device heterogeneity) and outline next steps, including micro-randomized trials for causal mechanism estimation, fairness audits, and longer-term retention studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Real-time cognitive and emotional state tracking in intelligent tutoring systems for enhanced learning outcomes

  • Kai Kong,
  • Haytham F. Isleem,
  • Rajanikanth Aluvalu,
  • Ghanshyam G. Tejani,
  • Ahmed Sayed M. Metwally

摘要

Background

Adapting instruction to learners’ moment-to-moment cognitive load and achievement emotions remains challenging in deployed Intelligent Tutoring Systems (ITS). We present a reliability-aware, real-time multimodal framework that senses facial, vocal, physiological, and interaction signals, estimates joint affect–cognition states, and delivers state-contingent interventions under tight latency constraints.

Objective

To (i) design an end-to-end platform-agnostic pipeline that integrates Cognitive Load Theory and Control–Value accounts of achievement emotions within an Integrated Cognitive–Affective Learning Model; (ii) validate technical performance and educational impact in laboratory and classroom settings; and (iii) report effect sizes and uncertainty per reviewer guidance.

Methods

A randomized, mixed-methods evaluation contrasted four conditions (Full adaptive, Cognitive-only, Affective-only, Control). The fusion layer aligns and normalizes modalities (DTW • PCA) with reliability weighting. A temporal Transformer and a GRU estimate emotions and triarchic load. A constrained contextual bandit selects content, interface, metacognitive, or motivational actions. Participants (N = 271) completed math and science units over nine sessions spanning 3 weeks. Primary outcomes included pre- and post-learning assessments, as well as transfer. Process outcomes were measured in terms of time spent in engagement versus confusion, frustration, and boredom. UX and latency were also logged. Mixed-effects models yielded estimates with 95% CIs; sensitivity checks addressed sensor quality and site effects.

Results

The system met real-time targets (end-to-end latency < 200 ms; reaction time 520–820 ms) while maintaining an affect recognition accuracy of 82–88% (weighted F1 ≈ score 0.85). Adaptive instruction improved post-test performance compared to a strong static ITS baseline, with the most significant gains for mid-ability learners (approximately + 15% points, Cohen’s d approximately 0.8). Process measures showed higher time in Engagement (≈ 42% vs. 29%) and productive Confusion (≈ 18% vs. 12%), and lower Frustration/Boredom, with strong convergence between detected negative affect and self-reported frustration (r ≈ 0.7). Mechanism-level analyses indicated state-contingent efficacy: content adaptations during Confusion, interface adjustments during Frustration, metacognitive prompts during Engagement, and motivational supports during Boredom.

Conclusions

A theoretically grounded, reliability-aware multimodal controller can translate fine-grained state estimation into timely, pedagogically meaningful adaptations that scale across laboratory and classroom contexts. Benefits concentrate where scaffolding has the highest leverage (mid-ability band) and are mediated by improved regulation of cognitive load and affect. We report limitations (including more challenging anxiety detection, classroom noise, and device heterogeneity) and outline next steps, including micro-randomized trials for causal mechanism estimation, fairness audits, and longer-term retention studies.