<p>Sign language serves as the primary communication medium for individuals who are deaf or hard of hearing. Despite its critical importance, barriers persist in communication between the deaf community and the broader society, primarily due to limited sign language proficiency among the general population. While automated sign language recognition (ASLR) systems leveraging machine learning technologies offer a promising solution, existing approaches face challenges in optimizing the trade-off between computational efficiency and recognition accuracy. This study presents CrossViViT, a novel architecture that integrates cross-attention mechanisms with video vision Transformer networks to address these limitations. Drawing inspiration from multi-branch network architectures that combine diverse feature perspectives for flexible image recognition, our approach achieves both computational efficiency and high accuracy. The proposed model demonstrates exceptional performance on the Vietnamese Sign Language (VSL) dataset, achieving 92.47% accuracy in recognizing 50 distinct gestures across 8510 videos while maintaining computational efficiency at approximately 629 FLOPS.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-attention multi branch for Vietnamese sign language recognition: CrossViViT

  • Minh Hoang Chu,
  • Hoang Diep Nguyen,
  • Thi Ngoc Anh Nguyen,
  • Hoai Nam Vu

摘要

Sign language serves as the primary communication medium for individuals who are deaf or hard of hearing. Despite its critical importance, barriers persist in communication between the deaf community and the broader society, primarily due to limited sign language proficiency among the general population. While automated sign language recognition (ASLR) systems leveraging machine learning technologies offer a promising solution, existing approaches face challenges in optimizing the trade-off between computational efficiency and recognition accuracy. This study presents CrossViViT, a novel architecture that integrates cross-attention mechanisms with video vision Transformer networks to address these limitations. Drawing inspiration from multi-branch network architectures that combine diverse feature perspectives for flexible image recognition, our approach achieves both computational efficiency and high accuracy. The proposed model demonstrates exceptional performance on the Vietnamese Sign Language (VSL) dataset, achieving 92.47% accuracy in recognizing 50 distinct gestures across 8510 videos while maintaining computational efficiency at approximately 629 FLOPS.