<p>Image–text retrieval is a fundamental task in cross-modal learning, yet existing methods often struggle to capture fine-grained semantic relationships, limiting retrieval performance. To address this challenge, we propose MLFGF-CLIP, a cross-modal retrieval framework based on multi-level fine-grained feature fusion. Specifically, we design a multi-level adaptive fusion module that integrates complementary features from different network layers, effectively combining fine-grained visual details with high-level semantic representations. To further enhance discriminative capability, we introduce a self-distillation strategy that transfers global semantic knowledge to local representations, and a boundary hard negative mining mechanism that strengthens the model’s ability to distinguish semantically similar but mismatched pairs. We also construct a fine-grained Chinese dataset, NetProduct, consisting of over 7000 images and 2000 domain-specific product descriptions, to evaluate retrieval in specialized application scenarios. Extensive experiments on NetProduct, COCO-CN, and Flickr30k-CN demonstrate that our approach consistently improves retrieval accuracy, achieving notable gains over strong baselines. In particular, on NetProduct, MLFGF-CLIP surpasses the best baseline by 4.54% (text-to-image) and 5.00% (image-to-text) in Recall@1. These results confirm the effectiveness of our framework in advancing fine-grained cross-modal retrieval.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multilevel fine-grained feature fusion method for cross-modal image–text retrieval

  • Baolu Wu,
  • Hui Fang,
  • Zhenhua Shao,
  • Luyao Kang,
  • Ge Xu,
  • Tao Wang,
  • Yin Guan

摘要

Image–text retrieval is a fundamental task in cross-modal learning, yet existing methods often struggle to capture fine-grained semantic relationships, limiting retrieval performance. To address this challenge, we propose MLFGF-CLIP, a cross-modal retrieval framework based on multi-level fine-grained feature fusion. Specifically, we design a multi-level adaptive fusion module that integrates complementary features from different network layers, effectively combining fine-grained visual details with high-level semantic representations. To further enhance discriminative capability, we introduce a self-distillation strategy that transfers global semantic knowledge to local representations, and a boundary hard negative mining mechanism that strengthens the model’s ability to distinguish semantically similar but mismatched pairs. We also construct a fine-grained Chinese dataset, NetProduct, consisting of over 7000 images and 2000 domain-specific product descriptions, to evaluate retrieval in specialized application scenarios. Extensive experiments on NetProduct, COCO-CN, and Flickr30k-CN demonstrate that our approach consistently improves retrieval accuracy, achieving notable gains over strong baselines. In particular, on NetProduct, MLFGF-CLIP surpasses the best baseline by 4.54% (text-to-image) and 5.00% (image-to-text) in Recall@1. These results confirm the effectiveness of our framework in advancing fine-grained cross-modal retrieval.