Transformers revolutionized the field of deep learning which is traditionally dominated by CNN. Comparative to normal images, hyperspectral images are capable of capturing many sharp features in the spectral band of the electromagnetic spectrum which is a 3D spatiotemporal data. CNN is used for extracting local information but are constrained to a limited receptive field. To address this issue, transformers are capable of providing a powerful global representation, including the finer details. In this work, a Swin Transformer architecture and principal component analysis (PCA) is implemented for dimension reduction. Since it is a type of Vision transformer (ViT), it allows the comparison of the vehicular traffic images with that of the reference image. This is carried out by shifting the windows. The architecture incorporates multilayer perceptron model. Stability of the model is increased by normalization of the input/output data patterns. The developed model is evaluated using VisDrone dataset, meticulously compiled by the AISKYEYE team at Tianjin University. This algorithm effectively enhances the detection accuracy (98, 99, and 99%) of aerial images while ensuring real-time processing in a simultaneous manner. It resulted in good generalization capabilities and capturing minute features with increased accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Remote Sensing Image Classification Using Transformer-Based Approach for Monitoring Vehicular Traffic

  • G. Vijaya Gowri,
  • Prameeladevi Chillakuru,
  • B. Latha,
  • A. Ganesan,
  • V. Srividhya,
  • K. Sujatha,
  • N. Kanya,
  • N. P. G. Bhavani,
  • N. Kavitha

摘要

Transformers revolutionized the field of deep learning which is traditionally dominated by CNN. Comparative to normal images, hyperspectral images are capable of capturing many sharp features in the spectral band of the electromagnetic spectrum which is a 3D spatiotemporal data. CNN is used for extracting local information but are constrained to a limited receptive field. To address this issue, transformers are capable of providing a powerful global representation, including the finer details. In this work, a Swin Transformer architecture and principal component analysis (PCA) is implemented for dimension reduction. Since it is a type of Vision transformer (ViT), it allows the comparison of the vehicular traffic images with that of the reference image. This is carried out by shifting the windows. The architecture incorporates multilayer perceptron model. Stability of the model is increased by normalization of the input/output data patterns. The developed model is evaluated using VisDrone dataset, meticulously compiled by the AISKYEYE team at Tianjin University. This algorithm effectively enhances the detection accuracy (98, 99, and 99%) of aerial images while ensuring real-time processing in a simultaneous manner. It resulted in good generalization capabilities and capturing minute features with increased accuracy.