错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-Scale Motion Extraction for Enhanced Human Action Recognition Based on RGB and Skeleton Modalities

  • Ke Xu,
  • Hong Shen,
  • Chan Tong Lam

摘要

Multimodal Human Action Recognition networks based on 2D-CNN and temporal segmented sampling provide an effective approach for action modeling with low overhead. However, constrained by resource limitations, existing methods generally rely only on interactions between adjacent segments, resulting in poor performance for recognizing long-term and complex motions across multiple temporal segments. To overcome this limitation, we propose the Dual-Scale Motion Extraction (DSME) method and design distinct variants for the RGB and skeleton modalities. It establishes long-term dependencies through cross global segment modeling, while synergistically incorporating short-term motion details to enable fine-grained action recognition. Specifically, DSME models temporal features through two complementary branches: an adjacent segments differences branch that constructs temporal variations between consecutive segments to capture subtle motion; and an global segments comparison branch, which generates global semantic structures. Finally, channel compression is applied to branch fusion, mitigating the high computational complexity inherent in global temporal modeling. We validate our approach on the PKU-MMD dataset and demonstrate performance improvements of 0.5% (C-sub) and 0.3% (C-view) over benchmark methods with negligible computational cost increase.