Recently, transformer-based solutions have exhibited remarkable success in 3D human pose estimation (3D-HPE) by computing pairwise relations between joints. However, we observed that the conventional self-attention mechanism in 3D-HPE tends to overly focus on a tiny fraction of joints. Moreover, these overfocused joints often lack relevance to the performed actions, resulting in models that struggle to generalize across poses. In this paper, we address this issue by incorporating prior information on the human body structure through a plug-and-play Joint-Part Attention (JPA) module. Firstly, we design a Part-aware Weighted Aggregation (PWA) module to merge different joints into distinct parts. Secondly, we introduce a Joint-Part Cross-scale Attention (JPCA) module to encourage the model to attend to more joints. This is achieved by configuring joint tokens to query part tokens across two scales. In our experiments, we apply JPA to various transformer-based methods, demonstrating its superiority on Human3.6M, MPI-INF-3DHP, and HumanEva datasets. We will release our code.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation

  • Dengqing Yang,
  • Zhenhua Tang,
  • Jinmeng Wu,
  • Shuo Wang,
  • Lechao Cheng,
  • Yanbin Hao

摘要

Recently, transformer-based solutions have exhibited remarkable success in 3D human pose estimation (3D-HPE) by computing pairwise relations between joints. However, we observed that the conventional self-attention mechanism in 3D-HPE tends to overly focus on a tiny fraction of joints. Moreover, these overfocused joints often lack relevance to the performed actions, resulting in models that struggle to generalize across poses. In this paper, we address this issue by incorporating prior information on the human body structure through a plug-and-play Joint-Part Attention (JPA) module. Firstly, we design a Part-aware Weighted Aggregation (PWA) module to merge different joints into distinct parts. Secondly, we introduce a Joint-Part Cross-scale Attention (JPCA) module to encourage the model to attend to more joints. This is achieved by configuring joint tokens to query part tokens across two scales. In our experiments, we apply JPA to various transformer-based methods, demonstrating its superiority on Human3.6M, MPI-INF-3DHP, and HumanEva datasets. We will release our code.