错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual variational network for unsupervised cross-modal hashing

  • Xuran Deng,
  • Zhihang Liu,
  • Pandeng Li

摘要

Cross-modal retrieval is a natural and highly valuable need in the current multimedia content explosion era. This paper addresses the problem of unsupervised cross-modal hashing retrieval which enables efficient retrieval across different modalities (e.g., image-text) without class labels. Most previous methods try to align visual and text binary representations in the joint Hamming space, by independently learning encoding functions for respective modality domains. However, since the paired training data describes the same object from different modalities, one modality data exactly plays a complementary role in learning encoding function for the other modality data, which has been less explored. This paper presents a novel cross-modal retrieval framework, called deep dual variational hashing (DDVH), by exploring dual variational mappings between modalities to bridge the inherent modality gap. Specifically, DDVH consists of two sub-modules, which are visual variational mapping (VVM) and textual variational mapping (TVM). VVM generates semantic-preserved binary codes for visual modality samples via the Gaussian latent embeddings, and TVM learns visual-guided binary codes for the corresponding text modality data. These two sub-modules can be jointly optimized under the cyclic consistency mechanism. Such a dual variational mapping strategy enables DDVH to generate unified binary representations for two modalities by visual-semantic interaction in the Hamming space. Comprehensive experiments on three benchmarks demonstrate that our proposed DDVH approach yields significant improvements compared to the state-of-the-art methods.