Generative Adversarial Network Based Asymmetric Deep Cross-Modal Unsupervised Hashing
摘要
With the explosive growth of internet information, cross-modal retrieval has become an important and valuable frontier hotspot. Due to its low storage consumption and high search speed, deep hashing has achieved significant success in cross-modal retrieval. Current research on unsupervised cross-modal hashing algorithms mainly focuses on two aspects: extracting high-level semantic information from given instances’ raw data and designing network structures suitable for unsupervised learning. However, despite the abundance of unsupervised method research found in the literature, many current studies overlook the fact that the data distributions of different modalities are highly distinct. In fact, asymmetric network structures are more in line with cross-modal data learning. Therefore, this paper proposes an asymmetric deep cross-modal unsupervised hashing method based on generative adversarial networks (referred to as UDCMH-GAN algorithm). This method utilizes the image network channel as the reconstruction network to learn more valuable high-level semantic information, while the text network is set to a conventional network structure. The introduction of generative and adversarial mechanisms aims to achieve better modality fusion and bridge the semantic gap. The proposed method is validated on widely used datasets, and the results demonstrate that asymmetric learning methods are indeed more reasonable and accurate for different modalities.