<p>This study presents a robust and efficient framework for real-time sign language recognition and object tracking, designed to address the limitations of YOLO-based models in sign language recognition, such as complex background interference, multi-scale image variability, and computational inefficiencies. The proposed YOLO-Mamba model integrates the ODMamba backbone network, a Path Aggregation Feature Pyramid Network (PAFPN), and an optimized detection head to improve recognition accuracy, robustness, and computational efficiency. To enhance multimodal fusion, the framework incorporates a Local Fast Attention (LFA) module, which combines skeletal joint features extracted by OpenPose with image features obtained from YOLO-Mamba, thereby enriching the spatial and structural representation of gestures. Experimental evaluations conducted on the Hagrid dataset demonstrate that the proposed model outperforms traditional methods across diverse scenarios, delivering significant gains in accuracy and computational efficiency. The model is particularly adept at handling small targets, navigating complex backgrounds, and meeting real-time processing requirements. These findings underscore the scalability and effectiveness of the proposed framework in advancing gesture-based communication systems, while also broadening the scope of multimodal recognition techniques in human-computer interaction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research and application of sign language recognition and target tracking model based on YOLO-Mamba

  • Jialin Wu,
  • Tao Yang,
  • Jintao Meng

摘要

This study presents a robust and efficient framework for real-time sign language recognition and object tracking, designed to address the limitations of YOLO-based models in sign language recognition, such as complex background interference, multi-scale image variability, and computational inefficiencies. The proposed YOLO-Mamba model integrates the ODMamba backbone network, a Path Aggregation Feature Pyramid Network (PAFPN), and an optimized detection head to improve recognition accuracy, robustness, and computational efficiency. To enhance multimodal fusion, the framework incorporates a Local Fast Attention (LFA) module, which combines skeletal joint features extracted by OpenPose with image features obtained from YOLO-Mamba, thereby enriching the spatial and structural representation of gestures. Experimental evaluations conducted on the Hagrid dataset demonstrate that the proposed model outperforms traditional methods across diverse scenarios, delivering significant gains in accuracy and computational efficiency. The model is particularly adept at handling small targets, navigating complex backgrounds, and meeting real-time processing requirements. These findings underscore the scalability and effectiveness of the proposed framework in advancing gesture-based communication systems, while also broadening the scope of multimodal recognition techniques in human-computer interaction.