The ZS-SBIR task that aims to address the inherent heterogeneity problem between sketches and images, as well as the knowledge transfer problem. However, most existing works utilize models pretraining on a single modality, leading to an inability to effectively establish relationships between images and text, and preventing sufficient knowledge transfer. In this paper, we rethink and present a visual-language pretraining-driven model (termed VLPD) for ZS-SBIR task. Specifically, we leverage the CLIP model to extract features for sketches, images, and text, to achieve knowledge transfer. Meanwhile, we propose a semantic consistency enhancement method that exploits textual modalities to enhance semantic consistency between sketches and images, alleviating the inherent heterogeneity trouble. Extensive experiments on three ZS-SBIR datasets reveal that VLPD dramatically outperforms the SOTA single-modal pretrained models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual-Language Pretraining-Driven Zero-Shot Sketch-Based Image Retrieval

  • Qing Zhang,
  • Jing Zhang,
  • Feilong Bao,
  • Xiangdong Su,
  • Guanglai Gao

摘要

The ZS-SBIR task that aims to address the inherent heterogeneity problem between sketches and images, as well as the knowledge transfer problem. However, most existing works utilize models pretraining on a single modality, leading to an inability to effectively establish relationships between images and text, and preventing sufficient knowledge transfer. In this paper, we rethink and present a visual-language pretraining-driven model (termed VLPD) for ZS-SBIR task. Specifically, we leverage the CLIP model to extract features for sketches, images, and text, to achieve knowledge transfer. Meanwhile, we propose a semantic consistency enhancement method that exploits textual modalities to enhance semantic consistency between sketches and images, alleviating the inherent heterogeneity trouble. Extensive experiments on three ZS-SBIR datasets reveal that VLPD dramatically outperforms the SOTA single-modal pretrained models.