Abstract <p>Satellite image data classification is a crucial task for many applications such as urban planning, environmental monitoring, national border security, etc. In the era of artificial intelligence, neural network approaches for satellite image classification have shown good results. Transformer based approaches have completely transformed the artificial intelligence methods in the last five years. Initially the transformer approaches have been proposed for text processing. For computer vision problems, Vision Transformer has been proposed in year 2020, which is utilized for many research applications in areas of healthcare, satellite imagery, defence etc. Transformer based models have demonstrated that the attention mechanism plays a crucial role. In this paper, a specialized attention mechanism approach focused on spatial, spectral, and temporal features of satellite image combined with a vision transformer, is proposed. The proposed architecture is known as Context-Aware Vision transformer (CAViT) for satellite image classification. We applied the proposed model on publicly available three satellite image scene datasets i.e., University of California Merced (UCM) with 21 classes, Aerial Image Dataset (AID) with 30 classes, and Remote Image Scene Classification dataset of Northwestern Polytechnical University (NWPU-RESISC45) with 45 classes. We used performance metrics parameters as accuracy, recall, precision, F1-score, and confusion matrix to evaluate the model’s performance with different datasets. This model achieved an overall accuracy of 99.33% for UCM, 97.71% for AID, and 95.63% for NWPU-RESISC45 dataset. The model shows competitive results against other deep learning models. Our research paper revealed CAViT proficiency in the satellite image classification applications ranging from environment monitoring to urban planning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Context-Aware Vision Transformer for Satellite Image Classification

  • Himanshu Srivastava,
  • Anuj Kumar Bharti,
  • Akansha Singh

摘要

Abstract

Satellite image data classification is a crucial task for many applications such as urban planning, environmental monitoring, national border security, etc. In the era of artificial intelligence, neural network approaches for satellite image classification have shown good results. Transformer based approaches have completely transformed the artificial intelligence methods in the last five years. Initially the transformer approaches have been proposed for text processing. For computer vision problems, Vision Transformer has been proposed in year 2020, which is utilized for many research applications in areas of healthcare, satellite imagery, defence etc. Transformer based models have demonstrated that the attention mechanism plays a crucial role. In this paper, a specialized attention mechanism approach focused on spatial, spectral, and temporal features of satellite image combined with a vision transformer, is proposed. The proposed architecture is known as Context-Aware Vision transformer (CAViT) for satellite image classification. We applied the proposed model on publicly available three satellite image scene datasets i.e., University of California Merced (UCM) with 21 classes, Aerial Image Dataset (AID) with 30 classes, and Remote Image Scene Classification dataset of Northwestern Polytechnical University (NWPU-RESISC45) with 45 classes. We used performance metrics parameters as accuracy, recall, precision, F1-score, and confusion matrix to evaluate the model’s performance with different datasets. This model achieved an overall accuracy of 99.33% for UCM, 97.71% for AID, and 95.63% for NWPU-RESISC45 dataset. The model shows competitive results against other deep learning models. Our research paper revealed CAViT proficiency in the satellite image classification applications ranging from environment monitoring to urban planning.