<p>Cross-Domain Zero-Shot Image Retrieval (CDZSIR) tackles retrieval tasks under unseen domains or categories. Recently, the vision-language models with prompt learning demonstrate strong representation capabilities. However, the deficiency in contextual semantic awareness and sub-optimal domain-invariant feature representation constrains robustness of these models in CDZSIR scenarios. To address this, we propose a dual-consistency learning framework integrating prompt reconstruction and semantic mining. This framework enhances semantic generalization by enforcing consistency at both visual and textual levels. Specifically, we introduce the Review-based Collaborative Semantic Prompter (RCSP), which utilizes a review mechanism and collaborative modules to mine textual semantics for consistency at the visual level. Furthermore, we propose the Semantic-Aware Reconstruction (SAR) module, which employs Variational AutoEncoder (VAE) to learn latent visual prompt representation. Assisted by text encoder, SAR can adaptively reconstruct domain-invariant visual prompt vectors, thereby ensuring visual-semantic consistency at the textual level. Extensive experiments show that our method exhibits considerable potential, delivering robust results on four challenging benchmark datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-consistency framework for cross-domain zero-shot image retrieval via prompt reconstruction and semantic mining

  • Haoxiang Zhang,
  • Rui Zhang,
  • He Jiang,
  • Deqiang Cheng,
  • Qiqi Kou

摘要

Cross-Domain Zero-Shot Image Retrieval (CDZSIR) tackles retrieval tasks under unseen domains or categories. Recently, the vision-language models with prompt learning demonstrate strong representation capabilities. However, the deficiency in contextual semantic awareness and sub-optimal domain-invariant feature representation constrains robustness of these models in CDZSIR scenarios. To address this, we propose a dual-consistency learning framework integrating prompt reconstruction and semantic mining. This framework enhances semantic generalization by enforcing consistency at both visual and textual levels. Specifically, we introduce the Review-based Collaborative Semantic Prompter (RCSP), which utilizes a review mechanism and collaborative modules to mine textual semantics for consistency at the visual level. Furthermore, we propose the Semantic-Aware Reconstruction (SAR) module, which employs Variational AutoEncoder (VAE) to learn latent visual prompt representation. Assisted by text encoder, SAR can adaptively reconstruct domain-invariant visual prompt vectors, thereby ensuring visual-semantic consistency at the textual level. Extensive experiments show that our method exhibits considerable potential, delivering robust results on four challenging benchmark datasets.