Advancing Arabic Sign Language Recognition: A Novel MobileNetv2-Based DL Framework with Superior Accuracy and Cross-Dataset Validation
摘要
Developing efficient and effective sign language recognition systems has drawn much interest as inclusive communication technologies grow in value. The subtle hand gestures, finger articulations, and dynamic stances of sign language hinder even achieving great accuracy and generalizability. This work develops the processing efficient and accuracy-boosting MobileNetv2 model to provide a deep learning-based solution to these difficulties. One can assess the proposed model’s performance, adaptability, and generalizability using top-tier models, ResNet-50, VGG-16, InceptionV3, EfficientNet, AlexNet, and Xception. The ArSL 2018 dataset is compared with these models. Furthermore, this study includes a cross-dataset evaluation conducted using the ArSL2L dataset. Furthermore, K-fold cross-valuation has been used to show the model’s generalizability; the model shows constant performance over all folds. With performance criteria ranging from 90 to 96%, results revealed that the model was far more resilient and versatile than rival models. The proposed model achieved a testing accuracy of 97.22%, demonstrating its effectiveness and robustness. To demonstrate the model’s real-time applicability, we evaluated latency and memory usage on a mobile device. The model achieved an average inference time of 42 ms per frame with a lightweight footprint of 13.6 MB, making it suitable for deployment on mobile platforms. Our study shows that the MobileNetv2 model performs well in identifying complicated gestures.