<p>Sign Language (SL) recognition is a crucial task in computer vision research to identify signs from the articulation of hand shapes, hand gestures, body movements and facial expressions. Building robust SL recognition systems requires large-scale datasets with accurate labels. However, such datasets are scarce for Indian Sign Language (ISL). This paper addresses this gap by introducing a large-scale dataset of isolated ISL signs. The dataset encompasses 2002 frequently used words in the deaf community, signed by 20 deaf adult signers and contains 40033 isolated videos. We split the dataset into train, validation and test sets with no signer intersection for benchmarking models. We follow the keypoint-based approach and report the results of some existing models on this dataset. Traditional architectures process spatial and temporal features sequentially. We introduce a novel keypoint-based model called the Hierarchical Windowed Graph Attention Transformer Encoder (HWGATE) that simultaneously aggregates spatial and temporal features using a unified spatio-temporal graph. Additionally, it enhances robustness through keypoint grouping. This model achieves an accuracy of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10044_2025_1529_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(95.67\%\)</EquationSource> </InlineEquation> on the presented dataset. We pre-trained the proposed model on the new dataset and then fine-tune it on the smaller ISL INCLUDE dataset, which yields state-of-the-art results of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10044_2025_1529_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="51" /> </InlineMediaObject> <EquationSource Format="TEX">\(98.28\%\)</EquationSource> </InlineEquation> accuracy, beating its baseline by 12.68 percentage points. The presented dataset and the model implementation code are available at <a href="https://cs.rkmvu.ac.in/%7eisl">https://cs.rkmvu.ac.in/~isl</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Windowed Graph Attention Transformer Encoder and a Large Scale Dataset for Indian Sign Language Recognition

  • Suvajit Patra,
  • Arkadip Maitra,
  • Megha Tiwari,
  • K. Kumaran,
  • Swathy Prabhu,
  • Swami Punyeshwarananda,
  • Soumitra Samanta

摘要

Sign Language (SL) recognition is a crucial task in computer vision research to identify signs from the articulation of hand shapes, hand gestures, body movements and facial expressions. Building robust SL recognition systems requires large-scale datasets with accurate labels. However, such datasets are scarce for Indian Sign Language (ISL). This paper addresses this gap by introducing a large-scale dataset of isolated ISL signs. The dataset encompasses 2002 frequently used words in the deaf community, signed by 20 deaf adult signers and contains 40033 isolated videos. We split the dataset into train, validation and test sets with no signer intersection for benchmarking models. We follow the keypoint-based approach and report the results of some existing models on this dataset. Traditional architectures process spatial and temporal features sequentially. We introduce a novel keypoint-based model called the Hierarchical Windowed Graph Attention Transformer Encoder (HWGATE) that simultaneously aggregates spatial and temporal features using a unified spatio-temporal graph. Additionally, it enhances robustness through keypoint grouping. This model achieves an accuracy of \(95.67\%\) on the presented dataset. We pre-trained the proposed model on the new dataset and then fine-tune it on the smaller ISL INCLUDE dataset, which yields state-of-the-art results of \(98.28\%\) accuracy, beating its baseline by 12.68 percentage points. The presented dataset and the model implementation code are available at https://cs.rkmvu.ac.in/~isl.