Synergistic spectral and spatial feature analysis with transformer and convolution networks for hyperspectral image classification
摘要
Hyperspectral imaging (HSI) contains several land cover objects with rich spatial and spectral features. By utilizing these features, deep convolution neural networks (CNN) improved HSI classification accuracy. However, shallow CNN lacks global co-relation of the spatial and spectral features. Further, by increasing the convolution layers, trainable parameters also increase. Hence, computation cost significantly increases. In this study, a fusion-based HFTNet model is designed that extracts features via convolution and transformer block to improve classification performance. In the proposed HFTNet, the convolution block extracts local semantic features, and the transformer block captures the attention-based global features. We reduced the computation costs by dividing the query vector into two parts and passing it to convolution and transformer blocks for feature extraction. Finally, features are combined to generate enhanced semantic local and global features. The effectiveness of the proposed method is tested on four datasets and achieved an accuracy of 99.34% (UP), 97.95% (IP), 99.70% (SV), and 84.23% (KSC). We found that HFTNet takes less computation time and achieves much better classification accuracy than other methods.