Spectral-Spatial Blockwise Masked Transformer With Contrastive Multi-View Learning for Hyperspectral Image Classification
摘要
Deep Learning methods have advanced in hyperspectral image (HSI) classification. However, acquiring high-quality labeled HSI data demands substantial human resources. Moreover, the correlation between spatial and spectral features may cause target confusion and computational challenges. To address these issues, we propose a pretrain and few-shot finetune framework, spectral-spatial blockwise masked transformer with contrastive multi-view learning (SS-MTC). The HSI cube is transformed into spectral-spatial tokens through blockwise patch embedding. Following it, the spatial-spectral encoder extracts dimensional features and utilizes spectral-spatial associate positional encoding to capture dimensional correlations. Masked reconstruction is achieved by constructing a masked label recovery task using a block-level random masking approach and obtaining the mask reconstruction loss. Contrastive multi-view learning is employed to learn more discriminative feature representations across different views of the same sample, thereby obtaining contrastive loss. The above two losses are weighted and combined as the total loss for pretraining. Then, the encoder is retained, and a classifier is added to finetune model parameters only using a small number of samples. Experimental results on Indian Pines (IP), Houston (HU), and Pavia University (PU) datasets demonstrate that SS-MTC achieves higher classification accuracies compared to other methods.