<p>The aging population has elevated falls into a critical public health issue. While camera-based YOLO algorithms offer non-contact detection, standard YOLOv13 struggles with occlusion, similar postures, and high computational demands. To address this, we propose RDD-YOLO, a task-oriented architecture optimized for fall detection accuracy and efficiency. Rather than simply stacking existing modules, RDD-YOLO assigns RepViTBlock to backbone feature extraction, DySample to detail-preserving feature fusion, and DHead to multi-scale regression, so that each component plays a complementary role in the detection pipeline. Evaluated on refined URFD and MCF datasets, RDD-YOLO outperformed YOLOv13, RT-DETR, and Faster-RCNN. For the URFD and MCF datasets, the mAP reached 92.1% and 87.3%, respectively. Speed tests showed that on a laptop (in a WSL2 environment), inference speed increased from approximately 47 FPS to approximately 57 FPS. Through pruning and FP16 half-precision inference, the speed further increased to over 70 FPS, demonstrating the feasibility of this method in real-time, resource-constrained application scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RDD-YOLO: a real-time fall detection model incorporating a lightweight visual transformer with dynamic sampling

  • Haimin Luo,
  • Qiyu Zhong,
  • Qing Guo,
  • Xuelian Li

摘要

The aging population has elevated falls into a critical public health issue. While camera-based YOLO algorithms offer non-contact detection, standard YOLOv13 struggles with occlusion, similar postures, and high computational demands. To address this, we propose RDD-YOLO, a task-oriented architecture optimized for fall detection accuracy and efficiency. Rather than simply stacking existing modules, RDD-YOLO assigns RepViTBlock to backbone feature extraction, DySample to detail-preserving feature fusion, and DHead to multi-scale regression, so that each component plays a complementary role in the detection pipeline. Evaluated on refined URFD and MCF datasets, RDD-YOLO outperformed YOLOv13, RT-DETR, and Faster-RCNN. For the URFD and MCF datasets, the mAP reached 92.1% and 87.3%, respectively. Speed tests showed that on a laptop (in a WSL2 environment), inference speed increased from approximately 47 FPS to approximately 57 FPS. Through pruning and FP16 half-precision inference, the speed further increased to over 70 FPS, demonstrating the feasibility of this method in real-time, resource-constrained application scenarios.