Facial Keypoint-Based Segment-Level Driver Yawning Detection by Graph-Temporal Convolutional Neural Network Modeling
摘要
Driving while drowsy may lead to car accidents and other dangerous situations. Since yawning is an obvious sign of drowsiness, it is crucial to develop an effective method for its detection. A critical limitation of previous frame-level models is their ineffectiveness in distinguishing between actions with similar appearances, such as yawning and talking. Therefore, this paper introduces a new segment-level approach for real-time, high-accuracy yawning detection using our Graph-Temporal Convolutional Network (GTCN) model. The approach begins with facial keypoints detection in video clips using OpenPose, followed by yawning and other mouth behaviors detection through GTCN. Our model not only enhances spatial information processing through the incorporation of facial keypoints and supervised pre-training, but also integrates temporal models and comprehensive fine-tuning of the entire model, thereby effectively boosting its classification performance. Extensive experiments on public yawning detection datasets demonstrate GTCN's superiority. The model achieves state-of-the-art performance with 99.25% accuracy on YawDD, and 94.33% accuracy on Nthu-DDD under a zero-shot setting. Experiments also reveal that the GTCN model has good real-time performance in practice.