Sign language is a mode of communication that enables individuals with hearing or speech impairment, or both, to express themselves. Sign language recognition from videos has become a new challenge in this research field. This paper focuses on isolated sign language recognition, which involves recognizing and interpreting phrases or words expressed through gestures and hand movements through a short video. With the advancement of convolutional and recurrent neural network architectures in computer vision, this paper proposes efficient deep learning models to recognize American Sign Language (ASL). The models implemented in this paper are ResNet50, ResNet50 + BiLSTM, Xception, and Xception + BiLSTM. Overall, ResNet50 + BiLSTM performed the best, with training, validation, and test accuracies of 79.37%, 69.56%, and 52.17%, respectively. The models were trained and evaluated on a 10-gloss subset of the WLASL dataset. Furthermore, a comparative analysis was performed with the models proposed in other research papers implemented for the same purpose.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

American Sign Language Recognition Using Deep Learning Models

  • Gauri D. Revankar,
  • Smitha S. Kumar

摘要

Sign language is a mode of communication that enables individuals with hearing or speech impairment, or both, to express themselves. Sign language recognition from videos has become a new challenge in this research field. This paper focuses on isolated sign language recognition, which involves recognizing and interpreting phrases or words expressed through gestures and hand movements through a short video. With the advancement of convolutional and recurrent neural network architectures in computer vision, this paper proposes efficient deep learning models to recognize American Sign Language (ASL). The models implemented in this paper are ResNet50, ResNet50 + BiLSTM, Xception, and Xception + BiLSTM. Overall, ResNet50 + BiLSTM performed the best, with training, validation, and test accuracies of 79.37%, 69.56%, and 52.17%, respectively. The models were trained and evaluated on a 10-gloss subset of the WLASL dataset. Furthermore, a comparative analysis was performed with the models proposed in other research papers implemented for the same purpose.