Enhancing Satellite Image Classification Using Vision Transformers and Deep Learning Models
摘要
Since it provides a perceptive analysis of numerous environmental and geographic events, satellite image classification is a basic component of remote sensing. This work uses recent deep learning (DL) techniques, including Vision Transformers (ViT), to examine satellite imagery classification. Examining four models in great detail: Vision Transformers (ViT), VGG-16, CNN, and ResNet-50. CNN is here. There are 5631 high-resolution images in the dataset spanning four primary categories. The approach consists of data preparation, augmentation, and adding attention approaches to improve the model performance. With an incredible 99.62% accuracy in terms of testing, the experimental results reveal that the ViT model surpasses all the other models. It also reveals better memory, accuracy, and F-score. ResNet-50, VGG-16, and CNN models demonstrate a quite mediocre degree of performance with an accuracy range of 87–89%. The ViT method offers significant opportunities for use in remote sensing based on great accuracy and durability over earlier versions. The results show that ViT quite correctly picks out geographic correlations in satellite data. The exact and practical results of the research obtained from satellite images will help urban planning, disaster control, and environmental monitoring. Reliable picture categorization directs decisions in resource management, agriculture, and studies on climate change.