Dual-consistency framework for cross-domain zero-shot image retrieval via prompt reconstruction and semantic mining
摘要
Cross-Domain Zero-Shot Image Retrieval (CDZSIR) tackles retrieval tasks under unseen domains or categories. Recently, the vision-language models with prompt learning demonstrate strong representation capabilities. However, the deficiency in contextual semantic awareness and sub-optimal domain-invariant feature representation constrains robustness of these models in CDZSIR scenarios. To address this, we propose a dual-consistency learning framework integrating prompt reconstruction and semantic mining. This framework enhances semantic generalization by enforcing consistency at both visual and textual levels. Specifically, we introduce the Review-based Collaborative Semantic Prompter (RCSP), which utilizes a review mechanism and collaborative modules to mine textual semantics for consistency at the visual level. Furthermore, we propose the Semantic-Aware Reconstruction (SAR) module, which employs Variational AutoEncoder (VAE) to learn latent visual prompt representation. Assisted by text encoder, SAR can adaptively reconstruct domain-invariant visual prompt vectors, thereby ensuring visual-semantic consistency at the textual level. Extensive experiments show that our method exhibits considerable potential, delivering robust results on four challenging benchmark datasets.