Driving while drowsy may lead to car accidents and other dangerous situations. Since yawning is an obvious sign of drowsiness, it is crucial to develop an effective method for its detection. A critical limitation of previous frame-level models is their ineffectiveness in distinguishing between actions with similar appearances, such as yawning and talking. Therefore, this paper introduces a new segment-level approach for real-time, high-accuracy yawning detection using our Graph-Temporal Convolutional Network (GTCN) model. The approach begins with facial keypoints detection in video clips using OpenPose, followed by yawning and other mouth behaviors detection through GTCN. Our model not only enhances spatial information processing through the incorporation of facial keypoints and supervised pre-training, but also integrates temporal models and comprehensive fine-tuning of the entire model, thereby effectively boosting its classification performance. Extensive experiments on public yawning detection datasets demonstrate GTCN's superiority. The model achieves state-of-the-art performance with 99.25% accuracy on YawDD, and 94.33% accuracy on Nthu-DDD under a zero-shot setting. Experiments also reveal that the GTCN model has good real-time performance in practice.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial Keypoint-Based Segment-Level Driver Yawning Detection by Graph-Temporal Convolutional Neural Network Modeling

  • Kaihua Chen,
  • Tingting Zhu,
  • Shaofeng Li,
  • Yinxue Shi

摘要

Driving while drowsy may lead to car accidents and other dangerous situations. Since yawning is an obvious sign of drowsiness, it is crucial to develop an effective method for its detection. A critical limitation of previous frame-level models is their ineffectiveness in distinguishing between actions with similar appearances, such as yawning and talking. Therefore, this paper introduces a new segment-level approach for real-time, high-accuracy yawning detection using our Graph-Temporal Convolutional Network (GTCN) model. The approach begins with facial keypoints detection in video clips using OpenPose, followed by yawning and other mouth behaviors detection through GTCN. Our model not only enhances spatial information processing through the incorporation of facial keypoints and supervised pre-training, but also integrates temporal models and comprehensive fine-tuning of the entire model, thereby effectively boosting its classification performance. Extensive experiments on public yawning detection datasets demonstrate GTCN's superiority. The model achieves state-of-the-art performance with 99.25% accuracy on YawDD, and 94.33% accuracy on Nthu-DDD under a zero-shot setting. Experiments also reveal that the GTCN model has good real-time performance in practice.