<p>This paper introduces SignKeyNet, a novel multi-stream keypoint-based neural network designed to enhance lip-reading recognition and support foundational sign language understanding. The architecture decouples signer movements into three primary streams hands, face, and body using 133 pose keypoints extracted via pose estimation techniques. Each stream is processed independently using specialized attention modules, followed by an attention-based fusion mechanism that models cross-modal spatiotemporal dependencies. SignKeyNet is evaluated on the MIRACL-VC1 lip-reading dataset, achieving superior performance over baseline models such as HMMs, DTW, CNNs, LSTMs, and Two-stream ConvNets, with results including an accuracy of 0.85, a Word Error Rate (WER) of 0.12, and a Character Error Rate (CER) of 0.06. These results highlight the effectiveness of attention-driven, multi-modal architectures for visual speech recognition tasks. While the current evaluation focuses on lip-reading due to dataset constraints, the proposed architecture is extendable to full Sign Language Translation (SLT) systems. SignKeyNet demonstrates strong potential for real-time deployment in accessibility technologies, particularly for the deaf and hard-of-hearing communities.&#xa0;</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multistream attention based neural network for visual speech recognition and sign language understanding

  • Fatma M. Talaat,
  • Basma M. Hassan

摘要

This paper introduces SignKeyNet, a novel multi-stream keypoint-based neural network designed to enhance lip-reading recognition and support foundational sign language understanding. The architecture decouples signer movements into three primary streams hands, face, and body using 133 pose keypoints extracted via pose estimation techniques. Each stream is processed independently using specialized attention modules, followed by an attention-based fusion mechanism that models cross-modal spatiotemporal dependencies. SignKeyNet is evaluated on the MIRACL-VC1 lip-reading dataset, achieving superior performance over baseline models such as HMMs, DTW, CNNs, LSTMs, and Two-stream ConvNets, with results including an accuracy of 0.85, a Word Error Rate (WER) of 0.12, and a Character Error Rate (CER) of 0.06. These results highlight the effectiveness of attention-driven, multi-modal architectures for visual speech recognition tasks. While the current evaluation focuses on lip-reading due to dataset constraints, the proposed architecture is extendable to full Sign Language Translation (SLT) systems. SignKeyNet demonstrates strong potential for real-time deployment in accessibility technologies, particularly for the deaf and hard-of-hearing communities.