<p>Human action recognition based on 3D skeleton data remains a formidable challenge with significant implications in fields such as surveillance, human-computer interaction, and healthcare. Existing approaches predominantly focus on exploring the spatial relationships between joints to address action recognition. However, we argue that the temporal dynamics of joint movements are equally important for accurately capturing and representing human actions. In this work, we pioneer the exploration of frequency-domain information to enhance the model’s sensitivity to the global temporal dynamics of human actions. We introduce FTIReg, a novel model that seamlessly integrates features from both frequency and time domains, leading to more accurate action representations and improved recognition performance. Additionally, we propose two temporal invariant motion descriptors, which include velocity magnitude and the angle between velocities of consecutive frames, both of which are intrinsic to human actions and independent of joint positions, providing a robust characterization of action types that effectively captures the inherent differences between them. Extensive evaluations on the NTU RGB+D 60, NTU RGB+D 120 and Northwestern-UCLA datasets demonstrate the superior accuracy of our approach, highlighting its potential for advancing the state of the art in human action recognition. Our code will be made publicly available at <a href="https://anonymous.4open.science/r/FTIReg-7068">https://anonymous.4open.science/r/FTIReg-7068</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FTIReg: Frequency-time integration with invariant features for skeleton-based human action recognition

  • Yunze He,
  • Wei Jiang,
  • Qiangqiang Du,
  • Haoqiang Wang

摘要

Human action recognition based on 3D skeleton data remains a formidable challenge with significant implications in fields such as surveillance, human-computer interaction, and healthcare. Existing approaches predominantly focus on exploring the spatial relationships between joints to address action recognition. However, we argue that the temporal dynamics of joint movements are equally important for accurately capturing and representing human actions. In this work, we pioneer the exploration of frequency-domain information to enhance the model’s sensitivity to the global temporal dynamics of human actions. We introduce FTIReg, a novel model that seamlessly integrates features from both frequency and time domains, leading to more accurate action representations and improved recognition performance. Additionally, we propose two temporal invariant motion descriptors, which include velocity magnitude and the angle between velocities of consecutive frames, both of which are intrinsic to human actions and independent of joint positions, providing a robust characterization of action types that effectively captures the inherent differences between them. Extensive evaluations on the NTU RGB+D 60, NTU RGB+D 120 and Northwestern-UCLA datasets demonstrate the superior accuracy of our approach, highlighting its potential for advancing the state of the art in human action recognition. Our code will be made publicly available at https://anonymous.4open.science/r/FTIReg-7068.