MsF-HigherHRNet: Multi-scale Feature Fusion for Human Pose Estimation in Crowded Scenes
摘要
To solve the problems of occlusion and human scale variation in crowded crowd scenes, we propose a Multi-scale Fusion HigherHRNet (MsF-HigherHRNet) based on HigherHRNet, which integrates Residual Feature Augmentation (RFA) and Double Refinement Feature Pyramid Network (DRFPN). Firstly, the introduction of RFA further enriches the semantic information of the multi-scale feature maps. Secondly, with the help of spatial attention mechanism and deformable convolution design ideas, the DRFPN is proposed. When the feature maps are fed into the DRFPN, the occlusion problem in crowded crowd scenes is effectively solved. The experimental results on the CrowdPose dataset under the same experimental environment and image resolution show that the average accuracy of MsF-HigherHRNet is 69.7%, which is 1.7% higher than the average accuracy of HigherHRNet under the same configuration and has better robustness.