Human action recognition using ST-GCNs for blind accessible theatre performances
摘要
Audio descriptions present a tool that helps blind audience members assist theater performances by conveying visual information, such as actors’ gestures. However, its high production process cost and effort limit its availability. To address this, we propose a computer vision based system for automated actor gestures recognition, using the state-of-the-art spatio-temporal graph convolution networks (ST-GCNs) for skeleton-based action recognition via transfer learning technique. Hence, we evaluated the transferability of three pre-trained ST-GCNs: the first proposed spatio-temporal graph convolution network (ST-GCN), convolution network of two-stream adaptive graphs (2s-AGCN), and the multi-scale disentangled unified graph convolution network (MS-G3D). We used NTU-RGBD action benchmark as the source domain and collected a novel dataset: TS-RGBD, to serve as the target domain. We then proposed two configurations to accommodate the diversity between the source and target domains. Results showed that ST-GCNs exhibit positive transferability enhancing the models’ recognition performance in theatre contexts, promoting automated system for gesture accessibility in theaters.