Transformer-based architectures have achieved SOTA performance in monocular 3D human mesh recovery (HMR) but require high computational cost. It is observed that numerous mesh vertices impose a significant computational burden on the transformer architecture in high-dimensional scenarios and their connection with human joints further exacerbates this burden, which also complicates efforts to optimize computational efficiency. In this paper, we propose a transformer-based Joint and mesh Vertex Separation (JVS) framework to decouple the interactions between human joints and mesh vertices for efficient HMR. Specifically, the Low-Dim Mesh Attention (LDMA) is designed to specifically address the computational burden caused by separated mesh vertices in high-dimensional scenarios. For further enhancing efficiency, we introduce Human Topology Guided Pruning (HTGP), which leverages the separation of human joints to prune redundant visual tokens. Experiments on Human3.6M and 3DPW demonstrate that our model significantly improves efficiency while maintaining competitive performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topology-Aware Transformer for Efficient Human Mesh Recovery with Low-Dim Mesh Attention

  • Haoyu Lv,
  • Lili Chen

摘要

Transformer-based architectures have achieved SOTA performance in monocular 3D human mesh recovery (HMR) but require high computational cost. It is observed that numerous mesh vertices impose a significant computational burden on the transformer architecture in high-dimensional scenarios and their connection with human joints further exacerbates this burden, which also complicates efforts to optimize computational efficiency. In this paper, we propose a transformer-based Joint and mesh Vertex Separation (JVS) framework to decouple the interactions between human joints and mesh vertices for efficient HMR. Specifically, the Low-Dim Mesh Attention (LDMA) is designed to specifically address the computational burden caused by separated mesh vertices in high-dimensional scenarios. For further enhancing efficiency, we introduce Human Topology Guided Pruning (HTGP), which leverages the separation of human joints to prune redundant visual tokens. Experiments on Human3.6M and 3DPW demonstrate that our model significantly improves efficiency while maintaining competitive performance.