S \(^2\) JD: Segmented Skeleton Joint Descriptor for Human Action Recognition from 3D Posture Data
摘要
Human Action Recognition (HAR) has become a very interesting area of study because it can be used in so many different ways, such as for human-computer interaction, human-robot interaction, visual surveillance, and so on. However, the conventional color videos assisted HAR has limited performance as they are composed of different illuminations, viewpoints, clothing, etc. The emergence of 3D skeleton sequences sorts out these problems and shows a new direction for HAR. However, they share a susceptibility to motion and sound. Therefore, in this work, we advocate for a novel action descriptor, which we term the Segmented Skeleton Joint Descriptor (S \(^2\) JD). S \(^2\) JD encodes the Spatio-temporal movements of an action and ensures improved recognition accuracy especially for actions with similar movements. Initially, S2JD segments each frame of the skeleton sequence into different local segments, and each segment is encoded with the distances between local joints. Next, S \(^2\) JD encodes all the frames which are separated by different time instances. After extensive simulation of the proposed method on different standard datasets, it gained an average accuracy of 76.1% which shows its superiority. It shows an average improvement of 3.2% from the existing methods.