This study presents a novel approach for enhancing American Sign Language (ASL) recognition using Graph Convolutional Networks (GCNs) integrated with successive residual connections. The proposed method leverages the MediaPipe framework to extract 21 key landmarks from each hand gesture, which are then used to construct graph representations. We introduce a robust preprocessing pipeline that includes translational and scale normalization techniques to ensure consistency across the dataset. The constructed graphs are subsequently fed into a GCN-based neural architecture, which employs residual connections to mitigate gradient-related issues and improve network stability. The architecture’s performance was rigorously evaluated on the ASL Alphabet dataset, achieving state-of-the-art results with a validation accuracy of 99.14%. Extensive experimentation, including a 5-fold cross-validation, demonstrated the model’s superior generalization capabilities. The integration of dropout layers and batch normalization further enhanced the model’s robustness against overfitting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing ASL Recognition with GCNs and Successive Residual Connections

  • Ushnish Sarkar,
  • Archisman Chakraborti,
  • Tapas Samanta,
  • Sarbajit Pal,
  • Amitabha Das

摘要

This study presents a novel approach for enhancing American Sign Language (ASL) recognition using Graph Convolutional Networks (GCNs) integrated with successive residual connections. The proposed method leverages the MediaPipe framework to extract 21 key landmarks from each hand gesture, which are then used to construct graph representations. We introduce a robust preprocessing pipeline that includes translational and scale normalization techniques to ensure consistency across the dataset. The constructed graphs are subsequently fed into a GCN-based neural architecture, which employs residual connections to mitigate gradient-related issues and improve network stability. The architecture’s performance was rigorously evaluated on the ASL Alphabet dataset, achieving state-of-the-art results with a validation accuracy of 99.14%. Extensive experimentation, including a 5-fold cross-validation, demonstrated the model’s superior generalization capabilities. The integration of dropout layers and batch normalization further enhanced the model’s robustness against overfitting.