TWRLR: Composed Image Retrieval Based via Two-Way Reciprocity Learning and Reasoning
摘要
Composed Image Retrieval (CIR) seeks to effectively locate target images through multi-modal queries that integrate a reference image with a textual description of modifications. Current approaches mainly depend on one-way retrieval systems, which have difficulty in capturing the nuanced relationships between images and text across different modalities. To address this limitation, we propose a Two-Way Reciprocal Learning and Reasoning (TWRLR) framework, which establishes a closed-loop semantic alignment through bidirectional query mechanisms. In the forward path, TWRLR jointly embeds reference images and modification texts to retrieve target images. In the reverse path, it reconstructs reference images by generating inverse descriptive texts from target images, thereby enhancing bidirectional semantic correlations. Technically, we inject learnable directional tokens ([FW]/[BW]) into the BLIP text encoder to guide query directions and design a Multi-Head Attention Fusion module to dynamically align image regions with textual attributes. The experiments conducted on the Fashion-IQ and CIRR datasets reveal that the proposed TWRLR model exhibits a remarkable improvement in performance compared to baseline models.