Deep Spatiotemporal Network Based Indian Sign Language Recognition from Videos
摘要
The deaf community has substantial obstacles because of the communication barrier with hearing individuals. The traditional method of relying on sign language interpreters is not a cost-effective solution to address this issue. Existing systems for dynamic sign language recognition employ the CNN-LSTM framework, which has achieved reasonable performance. However, relying solely on spatial features extracted through CNN is inadequate for accurate recognition of sign language words. In this study, we propose a novel end-to-end deep spatiotemporal network for recognizing Indian sign language from videos. Our framework combines the extraction of deep spatial features using Inception-ResNet-V2 and the utilization of handcrafted spatiotemporal features obtained from the application of Volume Local Directional Number (VLDN). Furthermore, we introduce a new encoder-decoder network based on Long Short-Term Memory (LSTM) to effectively learn the spatiotemporal features. Lastly, we conduct a comprehensive experiment to demonstrate the performance of our proposed method.