This paper presents a novel fatigue detection method based on a video Transformer, specifically designed for detecting driver fatigue. Unlike traditional Transformers, the proposed method introduces innovations in feature extraction, attention mechanism design, and end-to-end architecture. Notably, it enhances the model’s understanding of overall image layout and fine-grained details by introducing conditional positional encoding, thereby improving the accuracy of fatigue feature extraction. Additionally, the method employs a factorized dot-product attention mechanism, effectively reducing computational complexity while ensuring robust temporal feature extraction. Furthermore, a feature scaling module is incorporated to more comprehensively perceive facial movements. The end-to-end deep learning architecture is well-suited to capturing the dynamic information relevant to fatigue detection, enhancing both efficiency and accuracy. The model’s generalization ability and stability have been validated through extensive cross-validation on various public datasets. In summary, the proposed method significantly enhances the accuracy and efficiency of fatigue detection, providing reliable technical support for the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Fatigue Detection Method Based on Video Transformer

  • Yuqing Zhong,
  • Tie Liu,
  • Yonghong Yang,
  • Zhuhong Shao,
  • Yuanyuan Shang,
  • Hui Ding

摘要

This paper presents a novel fatigue detection method based on a video Transformer, specifically designed for detecting driver fatigue. Unlike traditional Transformers, the proposed method introduces innovations in feature extraction, attention mechanism design, and end-to-end architecture. Notably, it enhances the model’s understanding of overall image layout and fine-grained details by introducing conditional positional encoding, thereby improving the accuracy of fatigue feature extraction. Additionally, the method employs a factorized dot-product attention mechanism, effectively reducing computational complexity while ensuring robust temporal feature extraction. Furthermore, a feature scaling module is incorporated to more comprehensively perceive facial movements. The end-to-end deep learning architecture is well-suited to capturing the dynamic information relevant to fatigue detection, enhancing both efficiency and accuracy. The model’s generalization ability and stability have been validated through extensive cross-validation on various public datasets. In summary, the proposed method significantly enhances the accuracy and efficiency of fatigue detection, providing reliable technical support for the field.