<p>Building on the outstanding performance of modern 2D pose estimation algorithms, the pipeline of inferring 3D human pose from detected 2D keypoints has gained widespread adoption in the field of 3D human pose estimation. To meet real-time demands in practical applications, we propose LiftMamba, a lightweight 3D pose estimation network built upon the Mamba architecture. The network comprises two modules: Temporal–Spatial Mamba and Temporal–Spatial Interaction. To address missing depth cues and keypoint occlusion that arise when lifting 2D poses to 3D, we propose Temporal–Spatial Mamba, a dual-stream module built on the Mamba architecture. This module enables keypoint sequences to interact efficiently across temporal and spatial dimensions, thereby modeling depth information and preserving coherence among joints. Furthermore, to reinforce relational dependencies among keypoint sequences, we introduce Temporal–Spatial Interaction, a multidimensional interaction module built on the Transformer framework. LiftMamba combines the Mamba architecture’s low computational complexity with the Transformer’s ability to model long-range token dependencies, achieving strong performance on the Human3.6M and MPI-INF-3DHP benchmarks. Relative to the similarly performing D3DP network, LiftMamba reduces parameter count by 35.5% and computational complexity by 99.1%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LiftMamba: Mamba-based lightweight network for 3D human pose estimation

  • Ma Li,
  • Dexiang Liu,
  • Xinguan Dai,
  • Hangbiao Gao

摘要

Building on the outstanding performance of modern 2D pose estimation algorithms, the pipeline of inferring 3D human pose from detected 2D keypoints has gained widespread adoption in the field of 3D human pose estimation. To meet real-time demands in practical applications, we propose LiftMamba, a lightweight 3D pose estimation network built upon the Mamba architecture. The network comprises two modules: Temporal–Spatial Mamba and Temporal–Spatial Interaction. To address missing depth cues and keypoint occlusion that arise when lifting 2D poses to 3D, we propose Temporal–Spatial Mamba, a dual-stream module built on the Mamba architecture. This module enables keypoint sequences to interact efficiently across temporal and spatial dimensions, thereby modeling depth information and preserving coherence among joints. Furthermore, to reinforce relational dependencies among keypoint sequences, we introduce Temporal–Spatial Interaction, a multidimensional interaction module built on the Transformer framework. LiftMamba combines the Mamba architecture’s low computational complexity with the Transformer’s ability to model long-range token dependencies, achieving strong performance on the Human3.6M and MPI-INF-3DHP benchmarks. Relative to the similarly performing D3DP network, LiftMamba reduces parameter count by 35.5% and computational complexity by 99.1%.