Sign language is crucial for communication among individuals with hearing or speech impairments. Automated recognition systems are essential for learning and translating different sign language variants. However, these systems often face high computational demands and large memory footprints, limiting their use in real-time, resource-constrained deployment on edge and embedded devices. This research develops an optimized pipeline for American Sign Language (ASL) recognition, comparing Binarized Neural Networks (BNNs) with traditional full-precision neural networks. Using Larq, a library for training binarized models, we leverage BNNs’ reduced memory and computational needs, suitable for embedded systems and edge devices. The study uses a dataset of ASL alphabet images, applying data augmentation to address data imbalance and occlusions. Experimental results show that while our best traditional model (ResNet50) achieves 96% accuracy, BNNs maintain competitive accuracy of up to 94% while reducing model size to as little as 4MiB. Critically, BNNs achieve nearly half the average inference time per image (approximately 4ms) compared to traditional models (approximately 8–9ms), demonstrating their suitability for real-time ASL recognition on embedded and edge devices without sacrificing performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing American Sign Language Recognition with Binarized Neural Networks: A Comparative Study with Traditional Models

  • Shakeef Ahmed Rakin,
  • Afif Alamgir,
  • Mohammed Intishar Rahman,
  • Md. Tahjid Ahsan,
  • Sifat Mahmud,
  • Md Tanzim Reza

摘要

Sign language is crucial for communication among individuals with hearing or speech impairments. Automated recognition systems are essential for learning and translating different sign language variants. However, these systems often face high computational demands and large memory footprints, limiting their use in real-time, resource-constrained deployment on edge and embedded devices. This research develops an optimized pipeline for American Sign Language (ASL) recognition, comparing Binarized Neural Networks (BNNs) with traditional full-precision neural networks. Using Larq, a library for training binarized models, we leverage BNNs’ reduced memory and computational needs, suitable for embedded systems and edge devices. The study uses a dataset of ASL alphabet images, applying data augmentation to address data imbalance and occlusions. Experimental results show that while our best traditional model (ResNet50) achieves 96% accuracy, BNNs maintain competitive accuracy of up to 94% while reducing model size to as little as 4MiB. Critically, BNNs achieve nearly half the average inference time per image (approximately 4ms) compared to traditional models (approximately 8–9ms), demonstrating their suitability for real-time ASL recognition on embedded and edge devices without sacrificing performance.