Human action recognition (HAR) is a crucial field in computer vision with applications ranging from video surveillance to human-computer interaction. This study explores an efficient framework for HAR by leveraging human keypoint extraction using Google’s Mediapipe and Long Short-Term Memory (LSTM) networks. The Mediapipe framework provides accurate and real-time human keypoint extraction, significantly reducing the computational complexity associated with raw video data processing. These keypoints, representing skeletal movements, are utilised as input to LSTM networks, which capture the temporal dependencies vital for action classification. The proposed method is evaluated on two benchmark datasets: UCF101 and Kinetics 400. UCF101 contains 13,320 video clips across 101 action classes, while Kinetics 400 features 400 human action categories. The combination of Mediapipe for feature extraction and LSTM for temporal modelling achieves 92.40% and 86.8% accuracy on UCF101 and Kinetics400, respectively. The results demonstrate the effectiveness of keypoint-based approaches for HAR. This study highlights the potential of lightweight, real-time human action recognition frameworks suitable for resource-constrained environments, especially for edgeAI and edge Robotics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human Action Recognition Using Mediapipe Holistic Keypoints: A Deep Learning Approach

  • Utkarsh Shandilya,
  • Vijeta Sharma,
  • Deepti Mishra

摘要

Human action recognition (HAR) is a crucial field in computer vision with applications ranging from video surveillance to human-computer interaction. This study explores an efficient framework for HAR by leveraging human keypoint extraction using Google’s Mediapipe and Long Short-Term Memory (LSTM) networks. The Mediapipe framework provides accurate and real-time human keypoint extraction, significantly reducing the computational complexity associated with raw video data processing. These keypoints, representing skeletal movements, are utilised as input to LSTM networks, which capture the temporal dependencies vital for action classification. The proposed method is evaluated on two benchmark datasets: UCF101 and Kinetics 400. UCF101 contains 13,320 video clips across 101 action classes, while Kinetics 400 features 400 human action categories. The combination of Mediapipe for feature extraction and LSTM for temporal modelling achieves 92.40% and 86.8% accuracy on UCF101 and Kinetics400, respectively. The results demonstrate the effectiveness of keypoint-based approaches for HAR. This study highlights the potential of lightweight, real-time human action recognition frameworks suitable for resource-constrained environments, especially for edgeAI and edge Robotics.