A Method for Vietnamese Sign Language Recognition Using a Fusion of RGB and Depth Information Based on an Inflated 3D ConvNet
摘要
We propose a method for Vietnamese Sign Language (VSL) recognition using a fusion of RGB and Depth information based on the Inflated 3D ConvNet (I3D) architecture. Our approach applies I3D feature extraction independently to RGB and Depth data, concatenates the extracted features, and processes them through the final classification layer. This method leverages the complementary strengths of both data types, enhancing recognition performance. Experimental results show that our method outperforms the original I3D and other fusion strategies, including Early Fusion and Late Fusion.