Modal Fusion-Enhanced Two-Stream Hashing Network for Cross Modal Retrieval
摘要
With the explosive growth of multimodal data, cross-modal hashing has emerged as a standout technique in the field of retrieval. Unsupervised hashing retrieval, which does not rely on image label information, has garnered widespread attention. However, existing unsupervised methods still face several common issues. Firstly, current methods often only consider either local or global single-feature extraction in image feature extraction. Secondly, the matrices generated by modal fusion lack sufficient discriminability and cannot effectively capture the feature differences between modalities. In this paper, we propose a Modal Fusion Enhanced Hashing Network (MFEH) for cross-modal retrieval. Firstly, we utilize CLIP and VGG for both coarse-grained and fine-grained feature extraction from images to obtain richer image features. Secondly, we propose an effective fusion method to construct fusion matrices between modalities. Subsequently, by adjusting the similarity weights of the fusion matrix between modalities, we shorten the distances between the most similar instance pairs and increase the distances between the most dissimilar instance pairs, thereby generating hash codes with higher discriminability. Extensive experiments demonstrate that MFEH achieves satisfactory precision retrieval performance on three commonly used datasets.