Research and application of sign language recognition and target tracking model based on YOLO-Mamba
摘要
This study presents a robust and efficient framework for real-time sign language recognition and object tracking, designed to address the limitations of YOLO-based models in sign language recognition, such as complex background interference, multi-scale image variability, and computational inefficiencies. The proposed YOLO-Mamba model integrates the ODMamba backbone network, a Path Aggregation Feature Pyramid Network (PAFPN), and an optimized detection head to improve recognition accuracy, robustness, and computational efficiency. To enhance multimodal fusion, the framework incorporates a Local Fast Attention (LFA) module, which combines skeletal joint features extracted by OpenPose with image features obtained from YOLO-Mamba, thereby enriching the spatial and structural representation of gestures. Experimental evaluations conducted on the Hagrid dataset demonstrate that the proposed model outperforms traditional methods across diverse scenarios, delivering significant gains in accuracy and computational efficiency. The model is particularly adept at handling small targets, navigating complex backgrounds, and meeting real-time processing requirements. These findings underscore the scalability and effectiveness of the proposed framework in advancing gesture-based communication systems, while also broadening the scope of multimodal recognition techniques in human-computer interaction.