Composed Image Retrieval (CIR) is a multimodal image retrieval task that locates target images using a reference image and a modification text. Current CIR methods employ contrastive learning to align query representations with target image representations, a process that requires large-scale high-quality triplets. However, the scarcity of manually annotated triplets poses a significant challenge. To address this issue, our study proposes a novel solution for generating accurate and high-quality triplets. First, we leverage a semantic segmentation model and a multimodal large language model to generate detailed textual descriptions of images. Then, we perform style learning on existing modification texts in the dataset to synthesize new modification texts with similar stylistic patterns. These new modification texts and images are combined to create new triplets, effectively augmenting the training data for CIR models. Experiments on FashionIQ and CIRR datasets using CIR models with CLIP, BLIP, and BLIP-2 backbones demonstrate that our generated triplets significantly improve model performance compared to baseline datasets and other triplet generation methods. These results validate the feasibility and effectiveness of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Triplet Generation Method for Composed Image Retrieval via Detailed Captioning

  • Rui Zhuo,
  • Jiali Bian

摘要

Composed Image Retrieval (CIR) is a multimodal image retrieval task that locates target images using a reference image and a modification text. Current CIR methods employ contrastive learning to align query representations with target image representations, a process that requires large-scale high-quality triplets. However, the scarcity of manually annotated triplets poses a significant challenge. To address this issue, our study proposes a novel solution for generating accurate and high-quality triplets. First, we leverage a semantic segmentation model and a multimodal large language model to generate detailed textual descriptions of images. Then, we perform style learning on existing modification texts in the dataset to synthesize new modification texts with similar stylistic patterns. These new modification texts and images are combined to create new triplets, effectively augmenting the training data for CIR models. Experiments on FashionIQ and CIRR datasets using CIR models with CLIP, BLIP, and BLIP-2 backbones demonstrate that our generated triplets significantly improve model performance compared to baseline datasets and other triplet generation methods. These results validate the feasibility and effectiveness of our proposed method.