<p>2D human pose estimation aims to localize human joints from an input image, where high-resolution representations are crucial for achieving accurate predictions. However, models that rely on such representations are typically computationally expensive, making them unsuitable for deployment on resource-constrained edge devices. Although existing lightweight methods reduce model complexity, they often incur noticeable performance degradation. To address these limitations, we propose SD-HRNet, a lightweight and efficient pose estimation framework that integrates a Spatial Information Grouping Module (SIGM) and a Structure-aware Attention Alignment Distillation (SAAD) strategy. Specifically, SIGM captures structured spatial relationships among joints by grouping spatial information and replacing the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(1 \times 1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1</mn> <mo>×</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation> convolutions in ShuffleNetV2 with lightweight G-Shuffle blocks, effectively reducing redundant computation while enhancing contextual modeling capability. To further mitigate the performance loss introduced by lightweight design, SAAD enriches the teacher network through a human-body mask segmentation task and then transfers structure-aware knowledge to the student via intermediate feature distillation and attention-based alignment, addressing feature dimension inconsistency and strengthening structural reasoning. Extensive experiments on the COCO, MPII, and CrowdPose datasets demonstrate that SD-HRNet achieves competitive accuracy while substantially reducing parameters and computational cost, highlighting its effectiveness for real-time human pose estimation on resource-limited platforms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SD-HRNet: lightweight human pose estimation via spatial grouping and attention alignment distillation

  • Huilin Liu,
  • Caiping Xiang,
  • Chenxi Hu,
  • Xinyue Wen,
  • Qiong Fang

摘要

2D human pose estimation aims to localize human joints from an input image, where high-resolution representations are crucial for achieving accurate predictions. However, models that rely on such representations are typically computationally expensive, making them unsuitable for deployment on resource-constrained edge devices. Although existing lightweight methods reduce model complexity, they often incur noticeable performance degradation. To address these limitations, we propose SD-HRNet, a lightweight and efficient pose estimation framework that integrates a Spatial Information Grouping Module (SIGM) and a Structure-aware Attention Alignment Distillation (SAAD) strategy. Specifically, SIGM captures structured spatial relationships among joints by grouping spatial information and replacing the \(1 \times 1\) 1 × 1 convolutions in ShuffleNetV2 with lightweight G-Shuffle blocks, effectively reducing redundant computation while enhancing contextual modeling capability. To further mitigate the performance loss introduced by lightweight design, SAAD enriches the teacher network through a human-body mask segmentation task and then transfers structure-aware knowledge to the student via intermediate feature distillation and attention-based alignment, addressing feature dimension inconsistency and strengthening structural reasoning. Extensive experiments on the COCO, MPII, and CrowdPose datasets demonstrate that SD-HRNet achieves competitive accuracy while substantially reducing parameters and computational cost, highlighting its effectiveness for real-time human pose estimation on resource-limited platforms.