Synergizing skeletal graph and deep features through transfer learning: a multimodal framework for action recognition
摘要
Human Action Recognition (HAR) is a cornerstone of intelligent computer vision systems, particularly in intelligent surveillance, healthcare monitoring, autonomous robotics and human computer interaction. Despite recent advances in deep learning, most existing approaches either rely solely on skeletal graph representations or exploit appearance-based features, thereby limiting the ability to capture the full spectrum of spatiotemporal dynamics. To address these challenges, we introduce a novel A multimodal framework for Action Recognition that tackles this issue by synergistically combining two influential modalities skeletal graph based features and deep learning based visual features. The extracted hybrid features are further enhanced through transfer learning, employing a fine-tuned ResNet-50 backbone to optimize feature discrimination across diverse action categories. Experiments are conducted on a curated subset of the UCF101 dataset, encompassing five complex action classes: Handstand Pushups, Handstand Walk, Pull Ups, Punch, and Push Ups. The multimodal fusion feature set is passed by a fully connected softmax classifier and gives the outcome of recognition accuracy of 96%. This helps significantly outperforming conventional single-stream methods. This approach consistently outperforms conventional single-modality models. This system is demonstrating robustness to variations in viewpoint, background complexity, and partial occlusions. The outcomes confirm that the complementary synergy between skeletal structure modeling and convolutional representations. It helps to enable a more robust and generalizable action recognition framework. This hybrid state-of-the-art multimodal system establishes a foundation for building reliable, scalable, and high-performance human action recognition systems across real-world applications.