<p>With the exponential growth of multimedia content, the demand for high-precision and scalable Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) systems has intensified. Traditional retrieval methods often struggle with the domain gap between abstract sketches and photorealistic images, while relying heavily on costly labeled datasets. This study proposes a novel three-stage framework that integrates self-supervised contrastive pre-training with a dual-branch attention-driven architecture. Our methodology leverages an incremental learning strategy and a distance-driven separability index to refine feature extraction with reduced reliance on manual supervision. By anchoring sketches, images, and textual descriptions into a shared semantic space via cross-modal alignment and employing an asymmetric deep hashing layer, the framework promotes both high-level semantic consistency and computational efficiency through binary Hamming distance matching. Extensive experiments on standard benchmarks demonstrate that the proposed method achieves state-of-the-art performance, reaching an mAP of 93.25 on the FGVC-Aircraft and 87.7 on the CUB-200-2011 datasets. These results demonstrate strong fine-grained retrieval performance and suggest potential for efficient large-scale deployment via binary Hamming distance matching.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Object image retrieval based on self-supervised learning and zero-shot learning

  • Alireza Nazari,
  • Kambiz Rahbar,
  • Hamidreza Moghassemi,
  • Ziaeddin Beheshtifard

摘要

With the exponential growth of multimedia content, the demand for high-precision and scalable Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) systems has intensified. Traditional retrieval methods often struggle with the domain gap between abstract sketches and photorealistic images, while relying heavily on costly labeled datasets. This study proposes a novel three-stage framework that integrates self-supervised contrastive pre-training with a dual-branch attention-driven architecture. Our methodology leverages an incremental learning strategy and a distance-driven separability index to refine feature extraction with reduced reliance on manual supervision. By anchoring sketches, images, and textual descriptions into a shared semantic space via cross-modal alignment and employing an asymmetric deep hashing layer, the framework promotes both high-level semantic consistency and computational efficiency through binary Hamming distance matching. Extensive experiments on standard benchmarks demonstrate that the proposed method achieves state-of-the-art performance, reaching an mAP of 93.25 on the FGVC-Aircraft and 87.7 on the CUB-200-2011 datasets. These results demonstrate strong fine-grained retrieval performance and suggest potential for efficient large-scale deployment via binary Hamming distance matching.