错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Hash Subspace from Large-Scale Multi-modal Pre-Training: A CLIP-Based Cross-modal Hashing Framework

  • Yuheng Ji,
  • Xingwei Zhang,
  • Gang Zhou,
  • Xiaolong Zheng,
  • Daniel Dajun Zeng

摘要

Multi-modal pre-training models act as a foundational approach for various multi-modal downstream tasks, including cross-modal retrieval. However, existing cross-modal retrieval models designed with multi-modal pre-training commonly rely on real-value storage, resulting in high computation and storage costs that hinder their applicability on large-scale practical applications. On the other hand, cross-modal retrieval models that utilize hash storage methods often overlook the potential benefits of multi-modal pre-training methods, leading to limited generalization and semantic extraction abilities for constructing a robust semantic subspace. Therefore, in this paper, we propose a cross-modal hashing framework called CCMH (CLIP-based Cross-Modal Hashing), which facilitates the transferability of a well-trained real-value semantic subspace to a hash semantic subspace. We conduct comparative analysis between CCMH and six commonly-used benchmarks under two datasets, and we find that our proposed CCMH shows competitive performance. The experimental results highlight the effectiveness of CCMH on transferring real-value semantic subspace to hash semantic subspace, and our method offers valuable insights for future research in the field of cross-modal hashing.