Enhancing Action Recognition Using 3D Skeleton Reconstruction: Integrating MotioNet with AcTv2
摘要
Human action recognition from skeleton data is one of the quickly growing research domains due to varied possible applications in fields such as security surveillance, human computer interaction, healthcare, entertainment applications, and others. In this paper, we implement a solution to improve the efficiency of the AcTv2 model which we have improved from the original AcT model. Since the input to AcT and AcTv2 models is in 2D data format, it restricts learning the depth and spatial information of three dimensions during the training of the model. As a result, it leads to incorrectness while recognizing actions having complex movement or occlusion of body parts. Therefore, a solution is proposed to improve the performance of AcTv2 by integrating MotioNet model for converting 2D skeleton data into 3D data as an input to AcTv2. The experimental datasets used in this study are KTH and UTD-MAH. The results demonstrate that use of 3D data significantly enhances recognition of complex actions, particularly when occlusion is present. However, for simple actions, no big gain is obtained by using 3D skeleton data instead of 2D. This study demonstrates that depth information has its benefit in improving the accuracy of models of action recognition.