Comparative Evaluation of Deep Learning Models for Image Captioning of Satellite Images
摘要
In our investigation, we present a thorough comparative examination of contemporary models aimed at generating descriptive narratives based on satellite image analysis, leveraging a fusion of methodologies such as computer vision (CV), natural language processing (NLP), and machine learning (ML). Employing an array of cutting-edge neural network architectures including VGG-19, ResNet-50, DenseNet-201, EfficientNet B7, UNet, and Inception V3, we meticulously analyze intricate features within satellite images to extract essential patterns and information. These features are subsequently integrated into a Long Short-Term Memory (LSTM) network, a robust variant of Recurrent Neural Network (RNN), thereby enhancing our capacity to efficiently interpret and analyze satellite image data. Through exhaustive experimentation on established benchmark datasets such as the Sydney Dataset and RSIC Dataset, we conduct a comprehensive comparative analysis to assess the efficacy of diverse models. The notable advancements achieved by our model in the realm of satellite image captioning hold significant promise for practical applications in domains such as image interpretation. The observed milestones represent a substantive progression within the field, showcasing adaptable utility across diverse sectors. Particularly noteworthy is the superior performance attained through the pairing of EfficientNet B7 (ENet) with LSTM, yielding accuracy rates exceeding 90% across both the Sydney Dataset and RSIC Dataset. This underscores the effectiveness of synergizing ENet with LSTM architectures for satellite image captioning tasks, thereby emphasizing its potential for real-world deployment.