Occlusion remains a core challenge in human pose estimation. Because visual evidence is missing, accurate localization of hidden keypoints must exploit both local anatomical cues and global pose consistency. Yet the ternary visibility labels that quantify occlusion severity in mainstream datasets are often ignored. Accordingly, we propose ViGTNet, a visibility-guided GCN–Transformer alternating architecture that explicitly incorporates these labels for efficient local–global collaboration. ViGTNet first predicts keypoint visibility, then strengthens feature propagation from fully visible to occluded keypoints along graph edges and suppresses occlusion noise within global self-attention, while additionally emphasizing partially occluded keypoints during training via a dynamic loss-reweighting scheme. ViGTNet improves overall AP by 1.0 on COCO and APHard by 0.8 on CrowdPose, demonstrating robust performance under occlusion.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visibility-Guided GCN-Transformer: Enhancing 2D Pose Estimation Under Occlusion

  • Yanyan Su,
  • Zhiliang Qiu,
  • Min Lu,
  • Jun Xiang,
  • Shenglian Lu

摘要

Occlusion remains a core challenge in human pose estimation. Because visual evidence is missing, accurate localization of hidden keypoints must exploit both local anatomical cues and global pose consistency. Yet the ternary visibility labels that quantify occlusion severity in mainstream datasets are often ignored. Accordingly, we propose ViGTNet, a visibility-guided GCN–Transformer alternating architecture that explicitly incorporates these labels for efficient local–global collaboration. ViGTNet first predicts keypoint visibility, then strengthens feature propagation from fully visible to occluded keypoints along graph edges and suppresses occlusion noise within global self-attention, while additionally emphasizing partially occluded keypoints during training via a dynamic loss-reweighting scheme. ViGTNet improves overall AP by 1.0 on COCO and APHard by 0.8 on CrowdPose, demonstrating robust performance under occlusion.