<p>This paper presents a novel dangerous behavior detection algorithm aimed at addressing challenges in construction worker action recognition, such as low accuracy, limited behavior categories, complex models, and detection errors. The proposed method integrates RGB data and skeleton information to improve detection performance. It enhances detection precision through EGFaster R-CNN, which incorporates a global channel spatial attention module into the feature pyramid network and leverages EfficientNet for a better balance between accuracy and efficiency. On this basis, DMHRNet introduces a novel multidimensional grouped attention mechanism and adopts a classification-based coordinate discretization strategy to improve localization robustness. The detected keypoints are then converted into spatiotemporal representations and processed by 3D CNNs to capture motion dynamics. Additionally, depthwise separable convolutions reduce computational cost, while optional fusion with RGB features enhances semantic understanding, resulting in a more comprehensive and discriminative representation of dangerous behaviors. Experiments were conducted on a self-built DB dataset, the COCO2017 dataset, and the PASCAL VOC 2012 dataset. Moreover, in comparison with other vision-based methods, the proposed method shows great advantage in reducing the number of parameters while still achieving competitive estimation precision. The optimized lightweight architecture facilitates efficient deployment on HPC platforms, ensuring real-time responsiveness in large-scale construction site monitoring while balancing computational efficiency and detection accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A dangerous behavior detection algorithm with the fusion of RGB data and skeleton information

  • Daojin Yao,
  • Hanxin Chen,
  • Zichen Yang,
  • Yan Chen,
  • Xiong Yin,
  • Wentao Dong,
  • Yongxiang Yu

摘要

This paper presents a novel dangerous behavior detection algorithm aimed at addressing challenges in construction worker action recognition, such as low accuracy, limited behavior categories, complex models, and detection errors. The proposed method integrates RGB data and skeleton information to improve detection performance. It enhances detection precision through EGFaster R-CNN, which incorporates a global channel spatial attention module into the feature pyramid network and leverages EfficientNet for a better balance between accuracy and efficiency. On this basis, DMHRNet introduces a novel multidimensional grouped attention mechanism and adopts a classification-based coordinate discretization strategy to improve localization robustness. The detected keypoints are then converted into spatiotemporal representations and processed by 3D CNNs to capture motion dynamics. Additionally, depthwise separable convolutions reduce computational cost, while optional fusion with RGB features enhances semantic understanding, resulting in a more comprehensive and discriminative representation of dangerous behaviors. Experiments were conducted on a self-built DB dataset, the COCO2017 dataset, and the PASCAL VOC 2012 dataset. Moreover, in comparison with other vision-based methods, the proposed method shows great advantage in reducing the number of parameters while still achieving competitive estimation precision. The optimized lightweight architecture facilitates efficient deployment on HPC platforms, ensuring real-time responsiveness in large-scale construction site monitoring while balancing computational efficiency and detection accuracy.