Deep cross-modal hashing (DCMH) learning, with its powerful representation capabilities and superior efficiency, has established itself as an effective technique for rapid similarity search. However, most existing methods still struggle to fully capture the relationships of the multi-granularity hidden semantics of cross-modal and within-modal with overt semantics of the category information. To solve this problem, we propose a framework called contrastive multi-granularity hashing for cross-modal retrieval (CMGH). The framework hierarchically integrates coarse-grained and fine-grained conceptual features of different modalities into multi-granularity conceptual features. Specifically, we developed a interaction module of intra-modal that explores hidden semantic conceptual representations through multi-granularity concept learning. Additionally, we constructed a contrastive learning-based alignment module of inter-modal that synchronizes fine-grained conceptual representations. Moreover, we introduced a category hashing learning module, which leverages label co-occurrence enhancement to generate the hash code for all categories. These hash codes are further used to supervise the cross-modal hash codes generated by the cross-modal hashing learning module. Comprehensive experiments on several benchmark cross-modal datasets highlight the superior performance of CMGH.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMGH:Label Co-occurrence-Enhanced Contrastive Multi-granularity Hashing Cross-Modal Retrieval

  • Jiayi Yang,
  • Ziyong Lin,
  • Qinze Zhu,
  • Xiang Li,
  • Mingyong Li

摘要

Deep cross-modal hashing (DCMH) learning, with its powerful representation capabilities and superior efficiency, has established itself as an effective technique for rapid similarity search. However, most existing methods still struggle to fully capture the relationships of the multi-granularity hidden semantics of cross-modal and within-modal with overt semantics of the category information. To solve this problem, we propose a framework called contrastive multi-granularity hashing for cross-modal retrieval (CMGH). The framework hierarchically integrates coarse-grained and fine-grained conceptual features of different modalities into multi-granularity conceptual features. Specifically, we developed a interaction module of intra-modal that explores hidden semantic conceptual representations through multi-granularity concept learning. Additionally, we constructed a contrastive learning-based alignment module of inter-modal that synchronizes fine-grained conceptual representations. Moreover, we introduced a category hashing learning module, which leverages label co-occurrence enhancement to generate the hash code for all categories. These hash codes are further used to supervise the cross-modal hash codes generated by the cross-modal hashing learning module. Comprehensive experiments on several benchmark cross-modal datasets highlight the superior performance of CMGH.