Sign language recognition (SLR) plays a critical role in enabling seamless communication for individuals with hearing impairments. This study introduces a novel SLR framework leveraging vision transformers (ViTs) as the primary deep learning architecture. ViTs, renowned for their efficacy in image classification, are employed to process video frames for robust sign recognition. The framework includes a preprocessing pipeline to normalize input data, enhancing the quality of the learned representations. The proposed methodology is evaluated on two benchmark datasets: the Malaysian sign language dataset and the Chinese sign language dataset. Experimental results demonstrate the model’s capability, achieving accuracy rates of 92.39% and 90.26% on the Chinese and Malaysian datasets, respectively. These results underscore the potential of ViT-based architectures in recognizing signs across diverse linguistic and cultural contexts, paving the way for advanced, inclusive SLR systems. Furthermore, this approach highlights the scalability and adaptability of transformers for multimodal gesture recognition tasks, setting a foundation for future research in sign language translation and real-time systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViT Sign: An Effective Transformer-Based Approach for Sign Language Recognition

  • S. Renjith,
  • Aneesh Varghese,
  • S. S. Poorna,
  • K. Anuraj

摘要

Sign language recognition (SLR) plays a critical role in enabling seamless communication for individuals with hearing impairments. This study introduces a novel SLR framework leveraging vision transformers (ViTs) as the primary deep learning architecture. ViTs, renowned for their efficacy in image classification, are employed to process video frames for robust sign recognition. The framework includes a preprocessing pipeline to normalize input data, enhancing the quality of the learned representations. The proposed methodology is evaluated on two benchmark datasets: the Malaysian sign language dataset and the Chinese sign language dataset. Experimental results demonstrate the model’s capability, achieving accuracy rates of 92.39% and 90.26% on the Chinese and Malaysian datasets, respectively. These results underscore the potential of ViT-based architectures in recognizing signs across diverse linguistic and cultural contexts, paving the way for advanced, inclusive SLR systems. Furthermore, this approach highlights the scalability and adaptability of transformers for multimodal gesture recognition tasks, setting a foundation for future research in sign language translation and real-time systems.