<p>For deaf people and the hard-of-hearing population to communicate inclusively, Indian Sign Language recognition is essential. This work shows a strong vision-based pipeline that combines temporal modeling and keypoint-based feature extraction for categorizing solitary Indian Sign Language adjective gestures. After extracting two-dimensional hand and position landmarks from preprocessed gesture films using MediaPipe Holistic, the pipeline uses a proprietary Euclidean norm and interpolation approach for frame filtering and frame count normalization. A fixed-length representation of 30 keypoint frames per gesture is produced for efficient model input. Several Recurrent Neural Network architectures, such as Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (Bi-LSTM), Gated Recurrent Unit(GRU), and attention-enhanced variations, are trained and evaluated using the processed sequences. The novelty of this work lies in developing a real-time Indian Sign Language gesture recognition pipeline that integrates landmark-based representation, custom valid-frame filtering, Real-Time Intermediate Flow Estimation-based temporal interpolation, and gesture-preserving sequence resampling. It also introduces efficient keypoint-based augmentation and evaluates multiple Recurrent Neural Network variants for fine-grained adjective recognition—an area underexplored compared to American Sign Language or British Sign Language. Among the models, the Gated Recurrent Unit achieves the highest classification performance with an accuracy of 97.00%, macro F1-score of 0.97, and strong precision-recall balance, demonstrating its effectiveness in modeling temporal dynamics for gesture recognition. The proposed system holds promise for scalable, real-time sign language applications in educational and assistive technologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A landmark-based temporal pipeline for real-time adjective recognition in Indian sign language

  • Tanmaya Arora,
  • R. Girija,
  • Kesa Veera Venkata Yaswanth,
  • Pallavi Kumari

摘要

For deaf people and the hard-of-hearing population to communicate inclusively, Indian Sign Language recognition is essential. This work shows a strong vision-based pipeline that combines temporal modeling and keypoint-based feature extraction for categorizing solitary Indian Sign Language adjective gestures. After extracting two-dimensional hand and position landmarks from preprocessed gesture films using MediaPipe Holistic, the pipeline uses a proprietary Euclidean norm and interpolation approach for frame filtering and frame count normalization. A fixed-length representation of 30 keypoint frames per gesture is produced for efficient model input. Several Recurrent Neural Network architectures, such as Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (Bi-LSTM), Gated Recurrent Unit(GRU), and attention-enhanced variations, are trained and evaluated using the processed sequences. The novelty of this work lies in developing a real-time Indian Sign Language gesture recognition pipeline that integrates landmark-based representation, custom valid-frame filtering, Real-Time Intermediate Flow Estimation-based temporal interpolation, and gesture-preserving sequence resampling. It also introduces efficient keypoint-based augmentation and evaluates multiple Recurrent Neural Network variants for fine-grained adjective recognition—an area underexplored compared to American Sign Language or British Sign Language. Among the models, the Gated Recurrent Unit achieves the highest classification performance with an accuracy of 97.00%, macro F1-score of 0.97, and strong precision-recall balance, demonstrating its effectiveness in modeling temporal dynamics for gesture recognition. The proposed system holds promise for scalable, real-time sign language applications in educational and assistive technologies.