<p>Sign Language Recognition (SLR) involves identifying human actions that convey language, benefiting both deaf-mute individuals and facilitating interactions between humans and computers. SLR models capture linguistic features from upper body movements, which can be depicted as graphical representations. In each video frame, temporal and spatial information is extracted by understanding skeleton graphs and attention mechanism. This graph-based information will encompass both temporal and spatial semantics, enabling comprehension of sign language in videos. In this research, we propose a deep model, termed TeDG, to utilize the potential of graph-based representation using attention mechanisms. The graph is formed by extracting skeleton of the object’s upper body in video frames. Specifically, we employ prompt techniques to extract labels from sign language videos and then apply attention diffusion models and graph skeletons for recognition. Our experimental results demonstrate the effectiveness of TeDG compared to existing models on both our new dataset and widely public datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diffusion-guided graph convolutional networks for sign language recognition

  • Nam Vu Hoai,
  • Dat Tran-Anh

摘要

Sign Language Recognition (SLR) involves identifying human actions that convey language, benefiting both deaf-mute individuals and facilitating interactions between humans and computers. SLR models capture linguistic features from upper body movements, which can be depicted as graphical representations. In each video frame, temporal and spatial information is extracted by understanding skeleton graphs and attention mechanism. This graph-based information will encompass both temporal and spatial semantics, enabling comprehension of sign language in videos. In this research, we propose a deep model, termed TeDG, to utilize the potential of graph-based representation using attention mechanisms. The graph is formed by extracting skeleton of the object’s upper body in video frames. Specifically, we employ prompt techniques to extract labels from sign language videos and then apply attention diffusion models and graph skeletons for recognition. Our experimental results demonstrate the effectiveness of TeDG compared to existing models on both our new dataset and widely public datasets.