Dynamic Hierarchical Fusion of Foundation Features for Robust Pose Estimation
摘要
Precisely estimating the pose of novel objects in the presence of occlusions or lack of texture continues to pose a significant hurdle in computer vision research. Existing methods depend on high-quality reference images or manually designed features, but they struggle to maintain generalization in complex environments. This study aims to develop a method that can overcome the limitations of existing methods and improve the accuracy and robustness of object pose estimation for unseen objects in occlusion or texture-absent scenarios. This study proposes a dynamic feature fusion framework based on pre-trained foundational models, exploiting the complementary nature of hierarchical features from Stable-Diffusion and DINOv2. A multi-layer perceptron autonomously adjusts fusion weights, avoiding manual tuning. Experimental results demonstrate that the proposed method effectively reduces the pose estimation gap between unseen and seen objects in occluded scenes, improving it from −3.1% to +1.6%. The study confirms that the decoupling and dynamic fusion of hierarchical features in foundational models significantly enhance the robustness of pose estimation in complex scenes, providing an efficient solution for industrial inspection and robotic grasping.