An Intermediate Deep Feature Fusion Approach for Understanding Human Activities from Image Sequences
摘要
Human activity recognition (HAR) is complex in real time because of varying views, illuminations, backgrounds, and colors. With the current state of the art, deep learning (DL) algorithms are gaining more attention because of their automated feature extraction in contrast to the handcrafted machine learning (ML) methods. In this work, we aim to exploit a data fusion approach for HAR and propose an intermediate feature fusion approach for vision-based HAR employing convolutional neural networks (CNN2D) and transfer learning (TL) techniques with pretrained residual neural networks (ResNet50) for the extraction of local and global features, respectively. These extracted features are then fused employing a concatenation layer before classifying activities. We have focused on detecting two categories of activities: action (single person) and interactions (human–human and human-object). The proposed activity recognition is able to detect human activities in constrained as well as unconstrained environments with multiple viewpoints. The proposed work is evaluated with five benchmark vision datasets, namely, KTH, Weizmann, IXMAS, CASIA action database, and MSR Daily Activity 3D, in terms of accuracy and confusion matrix. This proposed framework is able to recognize complex activities with better accuracy than single-person-based activities seen in the MSR Daily Activity 3D and CASIA datasets, gaining the highest accuracy of 99.94% and 99.76%, respectively. The comparative analysis with the existing state-of-the-art methods shows the superiority of the performance of the proposed model in terms of accuracy.