Joint Modal Heterogeneous Balance Hashing for Unsupervised Cross-Modal Retrieval
摘要
Existing cross-modal hashing methods have made progress in enhancing retrieval capabilities and reducing model size, but they struggle to balance retrieval performance across different channels, leading to increased robustness.These methods often show low integration of multi-channel semantic information and fail to address image-text heterogeneity balance, focusing solely on enhancing retrieval accuracy, which can lead to high model robustness issues. We propose the Joint Modal Heterogeneous Balance Hashing for Unsupervised Cross-Modal Retrieval (JMBH) to address this. We utilise the large model CLIP to process raw data, facilitating multi-channel semantic integration. We then design multi-channel fusion modalities to explore co-occurrence information across channels and develop intra- and inter-channel constraints to mine this information. Extensive experiments on three datasets validate JMBH’s effectiveness in balancing image-text heterogeneity and reducing robustness.