Multimodal Collaborative Attention Fusion Network for Remote Sensing Visual Question Answering
摘要
Remote Sensing Visual Question Answering (RSVQA) aims to generate accurate answers to questions about remote sensing images. The inherent complexity of RS data, featuring rich local details and global contextual information, poses challenges in effectively modeling the deep interactions between these features. Additionally, significant cross-modal differences between remote sensing imagery and textual descriptions make it difficult for existing methods to capture deep semantic correlations across modalities. To address these issues, we propose the Multimodal Co-Attention Fusion Network (MCAF-Net), which incorporates a novel Dual-Stage Collaborative Attention Mechanism (DSCA) to enhance both intra-modal and cross-modal interactions. In the first stage, DSCA dynamically integrates local features extracted by a Convolutional Neural Network (CNN) with global features captured by a Vision Transformer (ViT), ensuring complementary feature synergy. In the second stage, DSCA performs deep fusion of the aggregated visual features with textual features extracted by BERT, capturing complex cross-modal dependencies. Extensive experiments on three public RSVQA datasets demonstrate that MCAF-Net outperforms state-of-the-art methods across all benchmarks, validating its effectiveness and robustness.