Multilingual speech to Indian sign language translation using synthetic animation: a resource-efficient approach
摘要
The technical development of sign language motion generation systems has great potential to make successful and effective communication between people. This study introduces an efficient system for translating multilingual spoken languages (English, Hindi, Punjabi) into Indian Sign Language (ISL) including non-manual features using a synthetic keyframe approach. A new Punjabi speech corpus is created, and speech recognition model is trained using the DeepSpeech toolkit. Three translation rules: direct, root word, and fingerspell translations are implemented to create intermediate notations for sign words. A multilingual corpus of 1036 words is constructed using the Hamburg Notation System, and gestures are rendered through a 3D signing avatar. The system achieved overall accuracy of 88% (English), 85% (Hindi), and 84% (Punjabi) for ISL. The Punjabi speech recognition model delivered high performance with a minimal Word Error Rate of 0.07% and 0.04% of Character Error Rate. Network tests performed and synthetic motion approach consumed minimal resources, requiring only 394.52 ms and 486 bytes, compared to video requests that took 1.33 s and 150 kilobytes. The synthetic animation approach has proven to be efficient for spoken to sign language translation, especially in scenarios where resources are critical and signing view space is important factor.