From coarse to fine: a two-stage common semantic space construction for unpaired cross modal retrieval
摘要
In recent years, a large number of cross-modal retrieval methods have been proposed and have achieved significant performance. However, these methods are based on large-scale image–text paired datasets, and collecting such datasets is costly and time-consuming. To deal with the problem, the task of unpaired cross-modal retrieval is proposed, in which paired annotation is unavailable during model training. We study this task and propose a two-stage common semantic space construction method, which is inspired by human process of learning about the world by first grasping things and then understanding the relationships between them. Firstly, we construct an entity common semantic space, in which only enable coarse retrieval due to the neglect of relationships between objects in scenes. We then extend the space to fusion common semantic space, focusing on the relationships between different regions in scenes, to more completely express objects in the scene and accurately calculate image–text similarity. We conducted comparisons, ablation studies and visualizations to evaluate the performance of our method, and further validate on a downstream task-noisy cross-modal retrieval, demonstrating significant improvements in accuracy to baseline methods and validating the effectiveness.