Enhanced Human Action Recognition from Synthetic UAVs Data Using Landmark Extraction and Mamba-Based Model
摘要
Human action recognition from UAVs has become essential. Numerous studies have demonstrated that large datasets, including synthetic data, are essential for training human action recognition models. In this context, we introduce an approach that leverages synthetic data to enhance performance and competitiveness in the field, which can effectively address the data shortage problem from traditional methods. This work combines Mamba, an effective framework, with MediaPipe’s pose estimation framework, significantly improving the extraction of skeletal key points from UAV footage. The process involves extracting crucial landmarks, calculating joint angles, and transforming these into a multivariate time series, which is then analyzed by a Mamba-based recognition model. The proposed method is tested on the RoCoG-v2 dataset with 57.5% accuracy - 17.3% higher than the state-of-the-art model. This result represents a significant advancement in human action recognition using UAVs, demonstrating the effectiveness of synthetic datasets in both research and real-world applications.