An Approach on Building Communicative Channel for Hand Sign Translation to Text and Speech Model
摘要
Recent breakthroughs in deep learning have yielded remarkable achievements in various video-related tasks. However, traditional deep learning models tend to emphasize only the most distinguishing features, often neglecting potentially valuable and nuanced content. This limitation severely hampers their ability to grasp the intricate visual grammars embedded in sign language videos, which rely on the synergy of diverse visual cues such as hand shapes, facial expressions, and body postures. In light of these challenges, the attempt to understanding video-based sign language hinges on the concept of multi-modal learning system. The presented system is developed by ensemble of multi-modal learning enabled DeepLabV3 architecture to handle the challenge of feature learning required both in the physical coordinates, with respect to dynamic time. Since the evaluation of gesture requires both physical coordinates as well as the variations in the positions with respect to time, these features are keenly monitored here with the help of neural computations. Various relativity between the changing coordinates is utilized to assure the gesture impressions. The proposed system considers an optimization model to collaborate the multiple modality changes dynamically generated through the video input. Segmented image regions are classified using DeepLabV3 computing blocks. The proposed DLV3 model achieved 98% accuracy.