错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic Preservation and Hash Fusion Network for Unsupervised Cross-Modal Retrieval

  • Xinsheng Shu,
  • Mingyong Li

摘要

With the exponential growth of multimedia content on the Internet, cross-modal retrieval has emerged as a critical research area. This task aims to efficiently and accurately retrieve relevant information across different modalities based on a given query. Despite some unsupervised cross-modal hashing methods proposed for image-text retrieval with unlabeled data, existing methods struggle with extracting semantic features, preserving multi-modal semantics, and ensuring modal interaction. To address these issues, we proposed the Semantic Preservation and Hash Fusion Network (SPHFN). Our approach includes a Long-Range Semantic Capturing module to capture extensive dependencies in text and an Identity Semantic Preservation module to maintain intrinsic semantic information of original samples. These representations are then combined to ensure semantic consistency across modalities. We also construct a joint similarity matrix to generate high-quality unified binary hash codes by leveraging the interaction among continuous codes from different modalities. Experimental results on multiple datasets show that our method significantly outperforms existing state-of-the-art approaches in cross-modal hashing retrieval tasks.