<p>Skeleton-based action recognition (SAR) using graph convolutional networks (GCNs) has become a prominent research area owing to its robustness against environmental interference. However, existing methods suffer from two critical limitations: adaptive adjacency matrices tend to neglect joint connectivity relationships during training, while asynchronous spatiotemporal feature extraction often leads to significant information loss. In this work, we propose a Local Spatiotemporal Feature Fusion Graph Convolutional Network (LSTF-GCN) for SAR. To address the limitations of adaptive adjacency matrix, we introduce a local aggregation adjacency matrix that constrains the learning process to accurately capture spatial dependencies among joints. For mitigating spatiotemporal information loss during feature extraction, a Spatio-Temporal Collaborative Fusion Attention is developed to dynamically allocate weights between spatial and temporal features, thus comprehensively modeling motion characteristics in both domains. Additionally, we design a Dual-stream Multi-scale Temporal Convolution module that constructs a two-branch architecture, embedding multi-scale temporal convolutions in parallel temporal feature extraction paths to enhance representation of action patterns across varying time spans. Extensive experiments on two large-scale datasets (NTU RGB+D and UAV-Human) validate that LSTF-GCN achieves superior performance in SAR, conclusively validating its advancement and the effectiveness of individual components.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LSTF-GCN: local spatio-temporal feature fusion graph convolutional network for skeleton-based action recognition

  • Runjie Li,
  • Ning He,
  • Chaoqun Wang,
  • Ruicheng Wang,
  • Wenhua Wang

摘要

Skeleton-based action recognition (SAR) using graph convolutional networks (GCNs) has become a prominent research area owing to its robustness against environmental interference. However, existing methods suffer from two critical limitations: adaptive adjacency matrices tend to neglect joint connectivity relationships during training, while asynchronous spatiotemporal feature extraction often leads to significant information loss. In this work, we propose a Local Spatiotemporal Feature Fusion Graph Convolutional Network (LSTF-GCN) for SAR. To address the limitations of adaptive adjacency matrix, we introduce a local aggregation adjacency matrix that constrains the learning process to accurately capture spatial dependencies among joints. For mitigating spatiotemporal information loss during feature extraction, a Spatio-Temporal Collaborative Fusion Attention is developed to dynamically allocate weights between spatial and temporal features, thus comprehensively modeling motion characteristics in both domains. Additionally, we design a Dual-stream Multi-scale Temporal Convolution module that constructs a two-branch architecture, embedding multi-scale temporal convolutions in parallel temporal feature extraction paths to enhance representation of action patterns across varying time spans. Extensive experiments on two large-scale datasets (NTU RGB+D and UAV-Human) validate that LSTF-GCN achieves superior performance in SAR, conclusively validating its advancement and the effectiveness of individual components.