Enhanced sign language translation using three stream multimodal occlusion resilient transformer
摘要
The proposed work addresses Continuous Sign Language Recognition and Translation (CSLRT) by introducing the Three-Stream Multimodal Occlusion-Resilient Transformer (T-MORT) model. T-MORT leverages three parallel encoder transformers to extract complementary features from visual, gesture, and emotion cues in sign language videos. These multi-modal representations are fused using a late fusion strategy, which preserves modality-specific information and enhances robustness to occlusions. A dual-decoder framework is employed: one decoder autoregressively generates gloss sequences to provide linguistic priors, while the second produces the final spoken language translation. To enhance efficiency, a dynamic keypoint extraction-based frame sampling algorithm, Spatio-Temporal Keypoint-Guided Sampling (STKGS), adaptively selects frames, reducing computational overhead. Extensive experiments on the PHOENIX14T dataset demonstrate state-of-the-art performance, surpassing existing CSLRT benchmarks. The proposed T-MORT model yields a BLEU-4 score of 25.07 and a ROUGE-L of 49.12.