Empowering classroom behavior recognition through hybrid spatial-temporal feature fusion
摘要
The recognition of learners’ behaviors in classrooms is an important problem. It supports the assessment of the level of interest and engagement of students in the lecture, helping instructors to adjust their teaching methods promptly. This paper proposes a novel approach to human behavior analysis in educational settings by leveraging a combination of the Swin Transformer and skeleton-based LSTM networks. By effectively processing both RGB video data and skeletal sequences, the proposed combination model outperforms other models in capturing intricate spatial-temporal features, leading to improved accuracy in detecting student behaviors such as reading, writing, sleeping, and raising hands. To facilitate this research, a comprehensive dataset is also constructed, combining RGB frames and skeleton structures to provide a rich representation of human actions. The experimental results demonstrate the effectiveness of the proposed method in accurately identifying student behaviors. This research serves as a foundation for developing advanced techniques to diagnose learners’ psychological states, paving the way for personalized and adaptive learning experiences.