A Review on Skeleton-Based Early Action Recognition
摘要
Skeleton-based early action recognition (EAR) has emerged as a crucial research area in computer vision, aiming to identify human actions from skeletal data at their early stages. We discuss the advantages of using skeleton data for EAR, including its action-specific nature, reduced computational cost, and robustness to occlusions. We also explore state-of-the-art feature extraction techniques, such as graph convolutional networks (GCNs), recurrent neural networks (RNNs), and transformers, and analyze their effectiveness in capturing temporal and spatial dependencies within skeletal data. Furthermore, we review various model architectures specifically designed for EAR, such as two-stream networks and hierarchical networks, and discuss their contributions to achieve accurate and efficient early action recognition. This review aims to provide a comprehensive understanding of current trends and advancements in skeleton-based EAR and stimulate further research in this promising area.