<p>With the growing adoption of skeleton-based action recognition in domains like smart monitoring and interactive human–machine systems, achieving low-latency prediction while maintaining high accuracy has become a critical challenge. This paper introduces a multi-scale online action prediction framework (MSOAP), which simultaneously models local joint dependencies and global motion patterns across spatial scales to continuously capture the dynamic evolution of skeleton sequences. To further enhance predictive performance, the framework incorporates a joint learning mechanism that optimizes action recognition together with future motion forecasting, while employing neural ordinary differential equations (neural ODE) to model the continuous temporal dynamics of skeleton sequences. These computations benefit from GPU-based parallel acceleration for low-latency online inference. This design not only improves the accuracy of predicting unobserved motion segments but also strengthens robustness for long-duration actions. Evaluations performed on standard skeleton datasets, such as NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA, indicate that the presented approach achieves competitive performance compared with representative baseline methods, particularly under low-observation online prediction settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MSOAP: multi-scale spatial skeleton representations for online action prediction

  • Hongwei Chen,
  • Jiaxin Guo,
  • Xinlu Zong,
  • Fangquan Cheng

摘要

With the growing adoption of skeleton-based action recognition in domains like smart monitoring and interactive human–machine systems, achieving low-latency prediction while maintaining high accuracy has become a critical challenge. This paper introduces a multi-scale online action prediction framework (MSOAP), which simultaneously models local joint dependencies and global motion patterns across spatial scales to continuously capture the dynamic evolution of skeleton sequences. To further enhance predictive performance, the framework incorporates a joint learning mechanism that optimizes action recognition together with future motion forecasting, while employing neural ordinary differential equations (neural ODE) to model the continuous temporal dynamics of skeleton sequences. These computations benefit from GPU-based parallel acceleration for low-latency online inference. This design not only improves the accuracy of predicting unobserved motion segments but also strengthens robustness for long-duration actions. Evaluations performed on standard skeleton datasets, such as NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA, indicate that the presented approach achieves competitive performance compared with representative baseline methods, particularly under low-observation online prediction settings.