ESD-Pose: Enhanced Semantic Discrimination for Generalizable 6D Pose Estimation
摘要
Existing generalizable object pose estimation frameworks utilize a set of reference images to predict the complete pose of the target object in a query scene, which does not require textured CAD models to generate training data and can handle unseen novel objects during inference. However, current methods suffer from insufficient discriminative capability due to the template matching strategy. Both potential distractors and negative samples with similar appearance can be confused with the foreground, which limits performance on precise pose estimation. To address these problems, we propose a novel method called ESD-Pose to enhance the discrimination capacity of the framework. Specifically, a semantic interaction aware (SIA) module is introduced to seek semantic consistency among reference images and discrepancies between reference-query pairs. This module mitigates problems related to model deception caused by distractors. For dealing with slender objects robustly, we propose a dynamic scale weight learner to generate adaptive weights for multi-scale feature fusion, making for reasonable utilization of semantic information at different levels. Finally, an IoU-guided loss is designed to align localization and scale prediction, thus facilitating accurate pose estimation. Comprehensive experiments in the LINEMOD and GenMOP datasets demonstrate that ESD-Pose outperforms existing advanced methods, further validating the effectiveness of our method.