Entity Semantic Feature Fusion Network for Remote Sensing Image-Text Retrieval
摘要
Recently, there has been remarkable progress in remote sensing image-text retrieval (RSITR), but in the past RSITR methods, researchers often try to extract features in images and texts from global and local perspectives, and the unique entity semantic contained in remote sensing images and texts rarely paid attention to, or even ignored. In this paper, we propose an Entity Semantic feature Fusion Network (ESFN), which uses the entity semantic in remote sensing images and texts to enhance the alignment degree and improve the retrieval accuracy. In the visual part, we propose a Scene Entity Filtering module (SEF), which can effectively extract significant entity semantic features from low-level feature maps. The Multi-level Adaptive Fusion module (MAF) adaptively selects the information of image features at different levels for feature fusion. In the textual part, we embed the entity semantic in the text into our textual feature extractor, so that it can have a good entity perception of remote sensing text. We designed a Text Phrase Enhancement module (TPE) to further extract and enhance entity semantic and alignment visual information in text. In addition, ESFN’s experimental results on RSICD and RSITMD datasets show that R@1 and meanRecall (mR) reach 8.14, 22.16, 18.81 and 37.70 respectively, which verifies the model’s perception of entity semantic in remote sensing images and texts. Through performance comparison, ablation study and visualization analysis, the effectiveness and superiority of this method are verified.