Image retrieval systems often struggle to align with a user’s search intent, particularly when the query image is composed of multiple semantic concepts. To address this challenge, we propose a novel tempered-weighting query-fusion approach that leverages the capabilities of pre-trained vision-language models (VLMs) to focus on specific image components expressed through textual descriptions. Notably, our method utilizes the BLIP-2 model’s embeddings without requiring additional training, enabling zero-shot Composed Image Retrieval (CoIR). For systematic evaluation of our approach and comparison, we introduce the Focus-CoIR dataset, a novel resource derived from the BDD100K dataset that adopts a focusing setting and provides multiple ground truth positive labels for each target image and corresponding textual description. Our experimental results on Focus-CoIR demonstrate the effectiveness of our approach, outperforming state-of-the-art CoIR methods across all evaluated ranking metrics. This work highlights the potential of pre-trained VLMs for CoIR and contributes to the advancement of CoIR research. We make the Focus-CoIR dataset and our implementation publicly available to support future research in this domain.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ZIRACLE: Zero-Shot Composed Image Retrieval with Advanced Component-Level Emphasis

  • Florian Erdösi,
  • Kilian Weishaupt,
  • Khanlian Chung

摘要

Image retrieval systems often struggle to align with a user’s search intent, particularly when the query image is composed of multiple semantic concepts. To address this challenge, we propose a novel tempered-weighting query-fusion approach that leverages the capabilities of pre-trained vision-language models (VLMs) to focus on specific image components expressed through textual descriptions. Notably, our method utilizes the BLIP-2 model’s embeddings without requiring additional training, enabling zero-shot Composed Image Retrieval (CoIR). For systematic evaluation of our approach and comparison, we introduce the Focus-CoIR dataset, a novel resource derived from the BDD100K dataset that adopts a focusing setting and provides multiple ground truth positive labels for each target image and corresponding textual description. Our experimental results on Focus-CoIR demonstrate the effectiveness of our approach, outperforming state-of-the-art CoIR methods across all evaluated ranking metrics. This work highlights the potential of pre-trained VLMs for CoIR and contributes to the advancement of CoIR research. We make the Focus-CoIR dataset and our implementation publicly available to support future research in this domain.