Spatio-temporal Representation Learning for Isolated Sign Language Recognition Using Transformer
摘要
Sign language recognition is a crucial tool for bridging communication gaps between the Deaf community and hearing individuals. This research presents a novel approach to isolated sign language recognition by leveraging spatio-temporal representation learning through transformer models. Our method addresses the limitations of traditional sequence models by capturing long-range dependencies and complex motion patterns inherent in sign language. We employ a holistic approach that integrates advanced data preprocessing techniques, including MediaPipe’s landmark extraction, to construct a robust dataset. The transformer model’s architecture is optimized for classifying signs with high accuracy, significantly outperforming existing models. Experimental results demonstrate the effectiveness of our approach, with detailed metrics. This research contributes to the development of real-time sign language recognition systems, offering potential applications in educational and assistive technologies.