<p>People with hearing impairments use Sign Languages (SLs) to communicate. They find it difficult to communicate with spoken-language users because spoken-language users do not understand SLs. We must encourage tools that allow sign language and spoken language users to communicate with one another. Sign language translation (SLT) attempts to translate sign-language videos into spoken language or vice versa. In India, the development of datasets for the Indian Sign Language (ISL) in India is progressing slowly due to researchers’ discrete efforts. Currently, there is no publicly available dataset on ISL to evaluate sentence-level Continuous Sign Language Translation (CSLT) approaches that can be used by a transformer-based model. In the proposed work, we present the first ISL Dataset for CSLT, ISH-NEWS, that contains 4,222 sentence videos of over 6.5K words. We use the transformer-based translation model to evaluate its performance against the ISH-NEWS dataset and establish a baseline for model performance. Using data augmentation techniques, we increased the proposed model’s BLEU-4 score by 8.46.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

End-to-end sentence-level Indian sign language translation with ISH-NEWS dataset and transformer model

  • Rina Damdoo,
  • Praveen Kumar,
  • Rahul Gogoi

摘要

People with hearing impairments use Sign Languages (SLs) to communicate. They find it difficult to communicate with spoken-language users because spoken-language users do not understand SLs. We must encourage tools that allow sign language and spoken language users to communicate with one another. Sign language translation (SLT) attempts to translate sign-language videos into spoken language or vice versa. In India, the development of datasets for the Indian Sign Language (ISL) in India is progressing slowly due to researchers’ discrete efforts. Currently, there is no publicly available dataset on ISL to evaluate sentence-level Continuous Sign Language Translation (CSLT) approaches that can be used by a transformer-based model. In the proposed work, we present the first ISL Dataset for CSLT, ISH-NEWS, that contains 4,222 sentence videos of over 6.5K words. We use the transformer-based translation model to evaluate its performance against the ISH-NEWS dataset and establish a baseline for model performance. Using data augmentation techniques, we increased the proposed model’s BLEU-4 score by 8.46.