From Screens to Peers: Using Deep Learning Models for Visual Attention Shifts Analysis in Technology-Enhanced Classrooms
摘要
Collaborative learning interactions in technology-enhanced learning environments are multi-modal, engaging students through verbal discussions, online interactions, body gestures, and gaze behaviors. Within Multi-Modal Learning Analytics (MMLA), students’ gaze behaviors and visual attentions remain underexplored. Recent deep learning models in the AIED field enable automated visual attention detection—without dedicated hardware—toward diverse learning components such as peers, multiple screens, and worksheet in tech-rich classrooms. This study applies these deep learning models to examine students’ gazing behaviors and objects. Sixty-eight students (34 dyads) completed two one-hour collaborative tasks; 56 h of video were processed to generate moment-by-moment analysis of gaze objects, followed by sequential analysis of visual attention shifts across digital or non-digital components. Results show higher-performing groups shifting visual attention among their own screens, peers’ screens, and partner’s face, whereas lower-performing groups transition mainly between their own screen and off-task space. The findings reveal the different visual attention strategies applied by students in tech-rich classrooms, contributing to the current MMLA research regarding gaze behaviors. The study also highlights the potential of deep learning models for understanding students’ visual attention in everyday tech-rich classrooms.