DCC-DETR: a real-time lightweight gesture recognition network for home human–robot interaction
摘要
Gesture recognition, an intuitive and efficient human–robot interaction method, shows great potential in smart living applications. However, deployment in home environments faces challenges from complex backgrounds, hand-like distractors, and varying illumination, particularly for lightweight and real-time implementations. To address these limitations in balancing efficiency, speed, and accuracy, a novel lightweight static gesture recognition network, DCC-DETR, based on RT-DETR, is proposed. First, an improved StarNet backbone is developed to enhance feature extraction efficiency, optimizing performance without compromising accuracy. Second, a Cascaded Group Local Attention Network (CGLAN) is designed to improve the perception and processing of local information significantly. Third, a Context-Guided Spatial Reconstruction Feature Pyramid Network (CSRFPN) enhances multi-scale feature representation and robust feature integration. Finally, the Wise Focaler-ShapeIoU loss function enables robust bounding box regression, leading to improved localization precision. Extensive experiments on self-built, 100 Days of Hands, and EgoHands datasets demonstrate DCC-DETR’s superiority in detection accuracy, inference speed, and model efficiency, highlighting its practical utility across diverse scenarios. Compared to RT-DETR, DCC-DETR improves mAP50 by 1.0% and increases FPS by 105.8%, while reducing computation by 80.0%, parameters by 63.2%, and significantly compressing model size, making it highly suitable for resource-limited environments. When deployed on NVIDIA Jetson AGX Orin with TensorRT acceleration, DCC-DETR achieves an end-to-end GPU inference latency of 6.63 ms, 1.62