Real-time cognitive and emotional state tracking in intelligent tutoring systems for enhanced learning outcomes
摘要
Adapting instruction to learners’ moment-to-moment cognitive load and achievement emotions remains challenging in deployed Intelligent Tutoring Systems (ITS). We present a reliability-aware, real-time multimodal framework that senses facial, vocal, physiological, and interaction signals, estimates joint affect–cognition states, and delivers state-contingent interventions under tight latency constraints.
ObjectiveTo (i) design an end-to-end platform-agnostic pipeline that integrates Cognitive Load Theory and Control–Value accounts of achievement emotions within an Integrated Cognitive–Affective Learning Model; (ii) validate technical performance and educational impact in laboratory and classroom settings; and (iii) report effect sizes and uncertainty per reviewer guidance.
MethodsA randomized, mixed-methods evaluation contrasted four conditions (Full adaptive, Cognitive-only, Affective-only, Control). The fusion layer aligns and normalizes modalities (DTW • PCA) with reliability weighting. A temporal Transformer and a GRU estimate emotions and triarchic load. A constrained contextual bandit selects content, interface, metacognitive, or motivational actions. Participants (N = 271) completed math and science units over nine sessions spanning 3 weeks. Primary outcomes included pre- and post-learning assessments, as well as transfer. Process outcomes were measured in terms of time spent in engagement versus confusion, frustration, and boredom. UX and latency were also logged. Mixed-effects models yielded estimates with 95% CIs; sensitivity checks addressed sensor quality and site effects.
ResultsThe system met real-time targets (end-to-end latency < 200 ms; reaction time 520–820 ms) while maintaining an affect recognition accuracy of 82–88% (weighted F1 ≈ score 0.85). Adaptive instruction improved post-test performance compared to a strong static ITS baseline, with the most significant gains for mid-ability learners (approximately + 15% points, Cohen’s d approximately 0.8). Process measures showed higher time in Engagement (≈ 42% vs. 29%) and productive Confusion (≈ 18% vs. 12%), and lower Frustration/Boredom, with strong convergence between detected negative affect and self-reported frustration (r ≈ 0.7). Mechanism-level analyses indicated state-contingent efficacy: content adaptations during Confusion, interface adjustments during Frustration, metacognitive prompts during Engagement, and motivational supports during Boredom.
ConclusionsA theoretically grounded, reliability-aware multimodal controller can translate fine-grained state estimation into timely, pedagogically meaningful adaptations that scale across laboratory and classroom contexts. Benefits concentrate where scaffolding has the highest leverage (mid-ability band) and are mediated by improved regulation of cognitive load and affect. We report limitations (including more challenging anxiety detection, classroom noise, and device heterogeneity) and outline next steps, including micro-randomized trials for causal mechanism estimation, fairness audits, and longer-term retention studies.