Domain generalization person re-identification (DG-ReID) aims to train a model on source data that can generalize well to an unseen target domain. Despite showing impressive performance, existing methods still struggle with performance degradation when source domain annotations are unavailable. To this end, this paper investigates domain generalization ReID in an unsupervised setting, where no labels are annotated for any source domains. In this work, we propose a novel and fast unsupervised domain generalization person ReID model, which can achieve the highest Rank-1 performance with only one epoch of training. Specifically, we introduce natural language supervision into the person ReID task, aiming to use pedestrian descriptions generated by a text generation model as supervision information. By using the contrast learning of pedestrian images and their corresponding descriptions, the image and text features of the same person can be closer together, avoiding the high time complexity of generating pseudo-labels. In addition, our framework does not require any labels for training (either real or pseudo labels) and thus can be easily applied to unsupervised person ReID, demonstrating competitive performance with respect to relevant methods. Extensive experiments validate the superiority of our method compared to existing state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text Based Unsupervised Domain Generalization Person Re-identification

  • Guoqing Zhang,
  • Tong Jin,
  • Tianqi Liu

摘要

Domain generalization person re-identification (DG-ReID) aims to train a model on source data that can generalize well to an unseen target domain. Despite showing impressive performance, existing methods still struggle with performance degradation when source domain annotations are unavailable. To this end, this paper investigates domain generalization ReID in an unsupervised setting, where no labels are annotated for any source domains. In this work, we propose a novel and fast unsupervised domain generalization person ReID model, which can achieve the highest Rank-1 performance with only one epoch of training. Specifically, we introduce natural language supervision into the person ReID task, aiming to use pedestrian descriptions generated by a text generation model as supervision information. By using the contrast learning of pedestrian images and their corresponding descriptions, the image and text features of the same person can be closer together, avoiding the high time complexity of generating pseudo-labels. In addition, our framework does not require any labels for training (either real or pseudo labels) and thus can be easily applied to unsupervised person ReID, demonstrating competitive performance with respect to relevant methods. Extensive experiments validate the superiority of our method compared to existing state-of-the-art methods.