Sign language translation(SLT) aims to directly translate sign language videos into corresponding spoken sentences. However, as a continuous video task, sign language translation still has some challenging problems: (a) the problem of extracting features with more recognizable regions and temporal information in video frames, and (b) the problem of semantic alignment between different modalities of video and text. To address these challenges, this paper proposes a network based on dual-stream module and iterative semantic-aware mechanism module. Specifically, we design an edge feature stream that can extract dynamic edge information of sign language to assist the global feature stream in utilizing highly recognizable edge information. Then, we design a multi-layer iterative semantic-aware mechanism to transfer the knowledge in the language model to the video model for enhancing the alignment of semantic information of the two modalities. Finally, we conduct experiments on the RWTH-PHOENIX-WEATHER-2014T(PHOENIX14T) and CSL-Daily datasets, and achieved a BLEU-1 effect of 49.66 on the PHOENIX14T dataset, which also shows that our method achieves significant translation effects and achieves a certain competitive performance on the SLT task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-Stream Iterative Semantic-Aware Mechanism for Sign Language Translation

  • Ailing Xia,
  • Jiaming Lu,
  • Yubao Chen,
  • Yuxi Zhou,
  • Jianhua Zhang

摘要

Sign language translation(SLT) aims to directly translate sign language videos into corresponding spoken sentences. However, as a continuous video task, sign language translation still has some challenging problems: (a) the problem of extracting features with more recognizable regions and temporal information in video frames, and (b) the problem of semantic alignment between different modalities of video and text. To address these challenges, this paper proposes a network based on dual-stream module and iterative semantic-aware mechanism module. Specifically, we design an edge feature stream that can extract dynamic edge information of sign language to assist the global feature stream in utilizing highly recognizable edge information. Then, we design a multi-layer iterative semantic-aware mechanism to transfer the knowledge in the language model to the video model for enhancing the alignment of semantic information of the two modalities. Finally, we conduct experiments on the RWTH-PHOENIX-WEATHER-2014T(PHOENIX14T) and CSL-Daily datasets, and achieved a BLEU-1 effect of 49.66 on the PHOENIX14T dataset, which also shows that our method achieves significant translation effects and achieves a certain competitive performance on the SLT task.