In the hearing-impaired community, sign language identification plays a critical role in breaking communication barriers. In this research, the authors used deep learning and machine learning models including Convolutional Neural Network (CNN), Visual Geometry Group Network (VGG), MediaPipe, Single-Shot Multibox Detector (SSD), and Autoencoders and Variational Autoencoders (VAEs) for Indian Sign Language Detection. The authors used multiple classification algorithms such as logistic regression, support vector machines, random forest classifiers, and multilayer perceptron to evaluate and compare the performance of these models, in terms of accuracy, precision, recall, and F1-score. Our goal for the paper was to identify the model that achieves the highest accuracy for the Indian Sign Language (ISL) recognition model. This research focuses on the comparative performance of different models, identifying the most efficient technique for achieving a reliable Indian Sign Language Recognition model. These outputs could enhance future research and development in the field of gesture-based communication systems, making them more accessible and reliable. After running the different models, the authors found out that the models based on SSD, VGG, and CNN have shown the best performance compared to others in the detection of Indian Sign Language. VGG model using random forest classifier attained an accuracy of 99.97%. The CNN model using logistic regression has achieved an accuracy of 99.96. This shows when VGG is integrated with random forest classifier, it excels in both. The SSD model with the random forest classifier achieved an accuracy of 99.94. The VGG model, with a random forest classifier, did well; the model produced high precision in the detection of the fine details that make for a more accurate classification. The two models of VGG-RF and CNN-logistic regression give very good accuracy. However, the VGG-RF model stands out as very powerful with ability to extract all the features from an image meaning the model makes the communication system of the hearing-impaired people work effectively as the model recognizes hand signs clearly and smooths the way through voice-activated systems or text-based interfaces. Such advancements will allow for real-time translation, hence a more accessible communication technology to close the accessibility gap of the hearing-impaired population. All models, SSD, CNN, and VGG, outperform other approaches by recognizing Indian Sign Language gestures, this model can support real-time translation tools that convert Indian Sign Language to written texts/language. Making it easier for communication between sign language users and the non-signers to understand each other very well. We can also integrate this model with video calls or phone calls so that it can be easier for specially abled people to communicate with others on different platforms. Our work tries to contribute to developing more inclusive, gesture-based communication, hoping to foster great independence in the hearing-impaired community.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Various Sign Language Detection Methods for Indian Sign Language Using Deep Learning

  • Shrawani Rayalkar,
  • Nayna Chavda,
  • Manan Sharma,
  • Kartik Nimhan,
  • Krish Bende,
  • Raounak Agarwala,
  • Harshvardhan Singh Rathaur

摘要

In the hearing-impaired community, sign language identification plays a critical role in breaking communication barriers. In this research, the authors used deep learning and machine learning models including Convolutional Neural Network (CNN), Visual Geometry Group Network (VGG), MediaPipe, Single-Shot Multibox Detector (SSD), and Autoencoders and Variational Autoencoders (VAEs) for Indian Sign Language Detection. The authors used multiple classification algorithms such as logistic regression, support vector machines, random forest classifiers, and multilayer perceptron to evaluate and compare the performance of these models, in terms of accuracy, precision, recall, and F1-score. Our goal for the paper was to identify the model that achieves the highest accuracy for the Indian Sign Language (ISL) recognition model. This research focuses on the comparative performance of different models, identifying the most efficient technique for achieving a reliable Indian Sign Language Recognition model. These outputs could enhance future research and development in the field of gesture-based communication systems, making them more accessible and reliable. After running the different models, the authors found out that the models based on SSD, VGG, and CNN have shown the best performance compared to others in the detection of Indian Sign Language. VGG model using random forest classifier attained an accuracy of 99.97%. The CNN model using logistic regression has achieved an accuracy of 99.96. This shows when VGG is integrated with random forest classifier, it excels in both. The SSD model with the random forest classifier achieved an accuracy of 99.94. The VGG model, with a random forest classifier, did well; the model produced high precision in the detection of the fine details that make for a more accurate classification. The two models of VGG-RF and CNN-logistic regression give very good accuracy. However, the VGG-RF model stands out as very powerful with ability to extract all the features from an image meaning the model makes the communication system of the hearing-impaired people work effectively as the model recognizes hand signs clearly and smooths the way through voice-activated systems or text-based interfaces. Such advancements will allow for real-time translation, hence a more accessible communication technology to close the accessibility gap of the hearing-impaired population. All models, SSD, CNN, and VGG, outperform other approaches by recognizing Indian Sign Language gestures, this model can support real-time translation tools that convert Indian Sign Language to written texts/language. Making it easier for communication between sign language users and the non-signers to understand each other very well. We can also integrate this model with video calls or phone calls so that it can be easier for specially abled people to communicate with others on different platforms. Our work tries to contribute to developing more inclusive, gesture-based communication, hoping to foster great independence in the hearing-impaired community.