Multi-scale Similarity Information Fusion Hashing for Unsupervised Cross-Modal Retrieval
摘要
Unsupervised cross-modal hashing has attracted more and more attention in the field of image and text retrieval due to its ability to operate without the constraints of label information. Although the single-scale similarity matrix has shown promising results in preserving multi-modal semantic correlations, it still fails to adequately represent the complex relationships between different modalities. Moreover, the presence of private information within modalities can have an adverse impact on the learning of unified hash codes. In this paper, we propose a novel Multi-scale Similarity Information Fusion Hashing (MSIFH) to address the above challenges. Specifically, we introduce a multi-scale similarity information fusion strategy by considering both global and local scale similarity matrices. This allows similarity information from different scales to complement each other and maintain the similarity relationship between the original data more comprehensively. Furthermore, we introduce a multi-level semantic aggregation module to adaptively extract high-level features for learning cross-modal shared semantics, avoiding interference from private information. Extensive experiments on two benchmark datasets show that the proposed MSIFH significantly outperforms the state-of-the-art baselines.