Semantic Information Enhancement and Semantic Alignment for Artistic Cross-Modal Retrieval
摘要
The main obstacle encountered in cross-modal retrieval centers around the semantic gap observed across different modalities. Existing approaches, which depend on pre-trained unimodal models and diverse cross-modal techniques, encounter challenges when attempting to grasp the high-level semantic associations required for efficient retrieval. To address this challenge, we propose an innovative framework that integrates semantic information enhancement with cross-modal semantic alignment technique. By leveraging co-occurrence relationships among concepts, our approach enhances the quality of vector representations and fosters a deeper comprehension of intrinsic semantic connections. Our cross-modal semantic alignment module enables meticulous alignment between heterogeneous modalities, further improving the retrieval performance. The proposed method achieves a remarkable performance on the public dataset, with a notable margin of 4.3% for R@1 accuracy compared to prior methods.