Combining Shape and Trajectory Features for Human Action Classification Using a Neural Network and Synthetic Data
摘要
Human activity recognition systems using visual content analysis algorithms use data collected from a variety of sensors, the most popular of which are RGB cameras. Highly accurate motion information can be recorded using motion capture and then used to generate synthetic human body models. The advantage of such data is the absence of other objects in the background, the visualization accuracy and the anonymity of a person. This paper proposes a modified action recognition approach which creates action representations using simple features observed over time, including shape measurements, ratios and centroid trajectory. A feature vector consists of shape descriptors calculated separately for each video frame. These are then normalised, transformed to the frequency domain and supplemented with trajectory information. Action representations are classified using a feed-forward neural network with one hidden layer and varying number of hidden neurons. The high effectiveness values obtained in the experiments show that the appropriate composition of elementary features of moving objects brings considerable benefits.