MS-STGCN: Multi-stream Spatio-temporal Graph Convolution Network for Activity Recognition Under Occlusion
摘要
Skeleton-based methods in Human Activity Recognition have gained significant attention due to their critical applications in analyzing activities and interpreting human behavior. Despite advancement in the field, challenges such as incomplete skeleton data, often caused by occlusion, remain unresolved. These issues significantly degrade the performance of existing models, which mainly depend on complete skeleton data. To address these limitations, we propose a framework using Multi-stream Spatial Temporal Graph Convolution Networks (STGCNs) to identify activities in partially occluded environments. The proposed method integrates pose encoding and multimodal feature generation techniques to extract discriminative features from joints, even in challenging scenarios. A key innovation is the activation-based multi-stream strategy, where under-activated key joints are passed to subsequent streams for further feature extraction, ensuring a comprehensive understanding of human activities. To accurately identify the activities and ensure resilience against occlusion effects, outputs from all streams are then combined and sent to a Long Short-Term Memory (LSTM). The extensive experimental results estimate the accuracy of the proposed model, achieving state-of-the-art performance on standard datasets and significantly qualifying the effects of occlusion in synthesized datasets.