MGAN: Attempting a Multimodal Graph Attention Network for Remote Sensing Cross-Modal Text-Image Retrieval
摘要
Cross-modal text image retrieval in remote sensing is a crucial task that requires the development of unified visual and textual representations. Previous research has primarily focused on global information or object features extracted through object detection algorithms to obtain local information. However, these studies have overlooked the complexity of remote sensing images, leading to insufficient utilization of local information. To address this issue, we propose the Multimodal Graph Attention Network (MGAN), which is based on visual graph neural networks. Our MGAN architecture includes a multi-level node information fusion module that utilizes different levels of object features to generate local information, compensate for the limitations of global information, and produce more expressive visual features. Additionally, we incorporate visual features into our model to guide the generation of text features, considering the correlation between text and objects in remote sensing images. We conduct extensive experiments on the RSITMD dataset, demonstrating that our method outperforms state-of-the-art methods by a margin of 2.27% in mR.