<p>Image retrieval is a fundamental task in computer vision and multimodal research, with critical applications in product search, person re-identification, and interactive robotics. Traditional methods typically rely on unimodal queries, which often fail to capture the complexity of user intent. Recent advancements have introduced composed image retrieval approaches that utilize multimodal queries to better align with user requirements. However, existing methods based on pre-trained vision-language models such as CLIP exhibit limitations: their embeddings lack geometric attributes and struggle to distinguish subtle image differences. Moreover, they insufficiently exploit the feature disparities between reference and target images and inadequately model relational variations in descriptive texts, resulting in suboptimal query-image matching. To address these challenges, we propose FCDG-CIR, a method that fine-tunes CLIP for Difference-Guided Composed Image Retrieval. This approach comprises two key strategies. First, it finetunes CLIP to align visual feature differences between reference and target images with textual descriptions, thereby improving differential reasoning capabilities. Second, it introduces a difference-guided multimodal feature adaptive fusion mechanism, employing a teacher-student framework to guide feature fusion based on both image and text differences. We evaluate our method on three benchmark datasets–FashionIQ, Shoes, and CIRR–with experimental results demonstrating its effectiveness and superior performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-tuning CLIP for difference-guided composed image retrieval

  • Cheng Xian,
  • Xiuyuan Li,
  • Mingyong Li

摘要

Image retrieval is a fundamental task in computer vision and multimodal research, with critical applications in product search, person re-identification, and interactive robotics. Traditional methods typically rely on unimodal queries, which often fail to capture the complexity of user intent. Recent advancements have introduced composed image retrieval approaches that utilize multimodal queries to better align with user requirements. However, existing methods based on pre-trained vision-language models such as CLIP exhibit limitations: their embeddings lack geometric attributes and struggle to distinguish subtle image differences. Moreover, they insufficiently exploit the feature disparities between reference and target images and inadequately model relational variations in descriptive texts, resulting in suboptimal query-image matching. To address these challenges, we propose FCDG-CIR, a method that fine-tunes CLIP for Difference-Guided Composed Image Retrieval. This approach comprises two key strategies. First, it finetunes CLIP to align visual feature differences between reference and target images with textual descriptions, thereby improving differential reasoning capabilities. Second, it introduces a difference-guided multimodal feature adaptive fusion mechanism, employing a teacher-student framework to guide feature fusion based on both image and text differences. We evaluate our method on three benchmark datasets–FashionIQ, Shoes, and CIRR–with experimental results demonstrating its effectiveness and superior performance.