<p>Researchers are rapidly turning their focus to human pose estimation as a crucial area of computer vision. In light of the shortcomings of existing Transformer-based pose estimate methods when handling localized features, this work presents MAQT, an enhanced end-to-end method aimed at precise multi-human body pose estimation. To improve the localization of keypoints that are sensitive to scale changes, MAQT offers an Asym-Fusion block. Additionally, we design a new query strategy to optimize the initial selection of queries with Uncertainty-minimal Query Selection. Two self-attention mechanisms are used in the decoding phase for understanding and recording spatial and semantic connections between keypoints. In this paper, the MAQT method is validated on the MS COCO and CrowdPose datasets, and favorable experimental results are obtained.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MAQT: multi-scale attention and query-optimized transformer for end-to-end pose estimation

  • Hong Liang,
  • Cuiping Wang,
  • Mingwen Shao,
  • Qian Zhang

摘要

Researchers are rapidly turning their focus to human pose estimation as a crucial area of computer vision. In light of the shortcomings of existing Transformer-based pose estimate methods when handling localized features, this work presents MAQT, an enhanced end-to-end method aimed at precise multi-human body pose estimation. To improve the localization of keypoints that are sensitive to scale changes, MAQT offers an Asym-Fusion block. Additionally, we design a new query strategy to optimize the initial selection of queries with Uncertainty-minimal Query Selection. Two self-attention mechanisms are used in the decoding phase for understanding and recording spatial and semantic connections between keypoints. In this paper, the MAQT method is validated on the MS COCO and CrowdPose datasets, and favorable experimental results are obtained.