错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sign Language Recognition (SLR): A Brisk Paired Deep Metric Attention Learning (BPDMAL) Model for Video Data Applications

  • P. V. V. Kishore,
  • D. Anil Kumar,
  • K. Srinivasa Rao

摘要

Existing sign language recognition (SLR) models have shown to lack precision in identifying a sign due to their inability to rationalize inter-class discriminations. Specifically, the trained SLR models are sensitive to small variations in hand movements and finger shapes across signs in a video sequence. To overcome the above problem, this work proposes to learn a class label by computing a metric variable that squeezes the displacement between within-class and across-class labels. Generally, metric learning is considerably slower than other deep SLR classification architectures on video data due to the triplet pairing process. In traditional triplet pairing, all frames in all classes participate during the training process in each episode. Contrastingly, this paper proposes a self-sourced singular pairing process between the anchor and positive frames along with an attention mechanism, resulting in Brisk Paired Deep Metric Learning (BPDMAL) model. The BPDMAL integrated with standard deep learning architectures is evaluated on our 2D video sign language dataset named KL2DSL and two other benchmark video-based sign language datasets. The proposed BPDMAL has improved performance over the traditional DML and state-of-the-art SLR Deep Learning Models with an incremental downfall in training and inferencing times making it a useful model for real-time deployment.