Semantic Prototypical Learning is Effective Guidance for Object Re-Identification
摘要
Object re-identification (Re-ID) is a critical task in computer vision that involves identifying and matching the same target across multiple camera views. Recent approaches often make additional use of task-oriented feature codec stacks in large network structures to extract useful information, such as appending the additional transformer attention block or designing the convolutional network structure using task-related multimodal features. However, these approaches often lead to more complex network designs as well as redundant features. Instead, we propose a feature-guided approach using semantic prototypes. We find that the inclusion of appropriately encoded lower-level semantic features can also directly guide the encoder to learn more efficient task-related information without the need for stacking higher-dimensional operators. To confirm this, we also design an edge-based low-level semantic extractor to provide guidance. Experimental results on the MSMT17, Market-1501, and VeRi-776 datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming existing baselines for both mAP and Rank-1 metrics. These results validate the effectiveness and versatility of our approach in tackling the challenges of Re-ID in diverse real-world environments.