错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Real-Time Word-Level Recognition of Indian Sign Language

  • R. Chaitanya Madhav,
  • Prasad Shet,
  • K. S. Mohan Murali,
  • R. Likhith,
  • H. R. Mamatha

摘要

Advancements in technology have catalyzed innovations in machine learning and artificial intelligence, particularly in image recognition. While existing research has predominantly focused on static sign recognition for Sign Language, this study proposes a real-time dynamic sign recognition of Indian Sign Language. Given the fundamental importance of communication, especially for the hearing and speech-impaired, this research addresses the critical need for instant translation of sign language into text. To resolve this communication gap, this research introduces a comprehensive dataset using references from the “Indian Sign Language Research and Training Centre New Delhi” an official government-backed YouTube Channel dedicated to Indian Sign Language, which standardizes Indian Sign Language across the country. The dataset comprises 60 signs encompassing verbs, animals, vegetables, legal terms, fruits, emotions and more. The authors utilized Mediapipe as a feature extractor to train a total of nine models, including seven deep learning models namely LSTM, CNN, GRU, Bidirectional LSTM, Stacked LSTM, Hybrid CNN-GRU, and Convolutional LSTM, along with two machine learning models, namely Random Forest and XGBoost. Notably, the Convolutional LSTM model achieved the highest accuracy of 99.88%. Reliable evaluation metrics like Mean Average Precision, Confusion Matrix, Accuracy, F1-Score were employed, which shows the correctness of the signs predicted and ensures the precision of dynamic sign recognition. The culmination of this research results in a real-time system capable of instantaneously identifying dynamic signs. This contribution spans diverse domains such as Pattern Recognition, Computer Vision, Video Recognition, etc., addressing key challenges like the need for gloves or additional sensors, the requirement for many videos to train a single sign, utilization of simple signs, high processing power demand, signer dependency, occlusions, pushing the boundaries of sign language recognition technology. Notable achievements include a reduction in inference time, real-time capability, adaptability to variable lighting conditions & robustness against occlusions. Additionally, the models demonstrate efficiency with smaller datasets, lesser training times, and independence from skin color of the signer. The research anticipates that this system will significantly contribute to bridging the communication gap for individuals with hearing or speech impairments. The integration of cutting- edge technology into this research addresses most of the research gaps and hence promises to significantly enhance the quality of life and communication for the hearing-impaired population, emphasizing the importance of addressing their unique needs.