Humans communicate to share information and build relationships with each other. This method uses modalities such as speech, hand gestures, and facial expressions to communicate with others. Despite the prevalence of spoken communication, individuals with speech and hearing impairments heavily rely on nonverbal modes of communication such as hand gestures and facial expressions. Hearing-impaired individuals communicate through visual language which uses hand gestures, facial expression, and body movements to convey the information. Indian Sign Language (ISL) is other sign language rich in syntax, context, and grammar. Education systems, other public sectors should have flexibility to communicate with hearing-impaired individuals as well. Hence this work aims to build Sign Spotting method which helps in the accessibility of signs in the sign language data. This work presents the method for spotting Indian sign language at word levels in a continuous sign language video. For this task, publicly available ISL data was used, but customized it as per requirement by annotating gloss signs in multiple Indian sign language videos which have starting and ending timestamp. Our work helps to predict the gloss signs and their respective start and end timestamps in a continuous ISL video. This work uses Inflated-3D model and Spatial–Temporal models for extracting important features from the videos. Then use a temporal module for concatenating and spotting the signs. The precision, recall, and F1-score of our work are 0.545, 0.514, and 0.529, respectively. This is calculated using Intersection over Union (IoU) technique to get a better insight about the results. This methodology aims to enhance the accuracy and adaptability of sign language recognition systems, contributing to a more effective and inclusive approach in interpreting Indian Sign Language.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sign Spotting for Indian Sign Language

  • Abdul Majid Mujahid,
  • F. M. Umadevi,
  • C. Sujatha

摘要

Humans communicate to share information and build relationships with each other. This method uses modalities such as speech, hand gestures, and facial expressions to communicate with others. Despite the prevalence of spoken communication, individuals with speech and hearing impairments heavily rely on nonverbal modes of communication such as hand gestures and facial expressions. Hearing-impaired individuals communicate through visual language which uses hand gestures, facial expression, and body movements to convey the information. Indian Sign Language (ISL) is other sign language rich in syntax, context, and grammar. Education systems, other public sectors should have flexibility to communicate with hearing-impaired individuals as well. Hence this work aims to build Sign Spotting method which helps in the accessibility of signs in the sign language data. This work presents the method for spotting Indian sign language at word levels in a continuous sign language video. For this task, publicly available ISL data was used, but customized it as per requirement by annotating gloss signs in multiple Indian sign language videos which have starting and ending timestamp. Our work helps to predict the gloss signs and their respective start and end timestamps in a continuous ISL video. This work uses Inflated-3D model and Spatial–Temporal models for extracting important features from the videos. Then use a temporal module for concatenating and spotting the signs. The precision, recall, and F1-score of our work are 0.545, 0.514, and 0.529, respectively. This is calculated using Intersection over Union (IoU) technique to get a better insight about the results. This methodology aims to enhance the accuracy and adaptability of sign language recognition systems, contributing to a more effective and inclusive approach in interpreting Indian Sign Language.