In the field of 3D motion analysis, enhancing anomaly detection through language-inspired occlusion-aware modeling is crucial. It can be effective in interpreting human motion under occlusion and aids the model in comprehending complex movement patterns. This paper introduces the OAD2D framework, which detects motion abnormalities by reconstructing 3D coordinates of mesh vertices and human joints from monocular videos, with a specific focus on occluded scenes. OAD2D utilizes optical flow to capture motion prior information from video streams, thereby enriching the data on occluded human movements and ensuring temporal-spatial alignment of poses. Furthermore, we innovate in abnormal behavior detection by integrating it with the Motion-to-Text model, which employs VQVAE to quantize 3D motion features. This method maps motion tokens to text tokens, facilitating a semantically interpretable analysis of motion and boosting the generalization of abnormal behavior detection through the use of a language model. Our approach demonstrates the robustness of anomaly detection against severe and self-occlusions, as it reconstructs human motion trajectories in global coordinates to effectively mitigate occlusion issues. We validate the effectiveness of our proposed method on the Human3.6M, 3DPW, and NTU RGB+D datasets, Outperforms other state-of-the-art methods by a substantial margin, with a 5.1% improvement in Accuracy and a 0.12 increase in \(F_1\) -Score. The code and data will be made publicly available upon publication.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Anomaly Detection in 3D Motion Through Language-Inspired Occlusion-Aware Modeling

  • Su Li,
  • Liang Wang,
  • Jianye Wang,
  • Ziheng Zhang,
  • Junjun Zhang,
  • Lei Zhang

摘要

In the field of 3D motion analysis, enhancing anomaly detection through language-inspired occlusion-aware modeling is crucial. It can be effective in interpreting human motion under occlusion and aids the model in comprehending complex movement patterns. This paper introduces the OAD2D framework, which detects motion abnormalities by reconstructing 3D coordinates of mesh vertices and human joints from monocular videos, with a specific focus on occluded scenes. OAD2D utilizes optical flow to capture motion prior information from video streams, thereby enriching the data on occluded human movements and ensuring temporal-spatial alignment of poses. Furthermore, we innovate in abnormal behavior detection by integrating it with the Motion-to-Text model, which employs VQVAE to quantize 3D motion features. This method maps motion tokens to text tokens, facilitating a semantically interpretable analysis of motion and boosting the generalization of abnormal behavior detection through the use of a language model. Our approach demonstrates the robustness of anomaly detection against severe and self-occlusions, as it reconstructs human motion trajectories in global coordinates to effectively mitigate occlusion issues. We validate the effectiveness of our proposed method on the Human3.6M, 3DPW, and NTU RGB+D datasets, Outperforms other state-of-the-art methods by a substantial margin, with a 5.1% improvement in Accuracy and a 0.12 increase in \(F_1\) -Score. The code and data will be made publicly available upon publication.