Cross-Modal Retrieval Based on Semantic Filtering and Adaptive Pooling
摘要
Studies on cross-modal retrieval have caught a lot of attention in recent years, in previous studies, researchers apply attention mechanisms to cross-modal retrieval tasks and gain significant performance improvement. However, the previous cross-attention mechanisms between visual and textual modals may lead to semantic misalignment issues. In this paper, we present Semantic Filtering and Adaptive Pooling to address the issue of semantic misalignment and improve the aggregation of local features into global features. Our approach outperforms multiple state-of-the-art models on the Flickr30K dataset.