A rotation-invariant horizontal vertical pooled module for remote sensing image representation
摘要
Accurate information retrieval from multi-source and multi-resolution image data constitutes a foundation for knowledge discovery. Scene image classification in the remote sensing (RS) community using aerial very high resolution (VHR) images is one of the well-researched areas, which mostly utilise deep learning (DL)—based methods thanks to their remarkable classification performance. Nevertheless, existing DL-based methods still have a limited ability to capture precise spatial semantic information scattered toward the horizontal and vertical directions across such images at multiple scales and rotations. As such, we herein propose a novel approach, employing an innovative rotation invariant horizontal vertical pooled module (RIHVPM), to well-represent aerial VHR RS images for stable and improved classification performance. Notably, the proposed RIHVPM benefits from the multiple tensor rotations coupled with attention-enabled multiscale horizontal and vertical pooling operations for image representation. An experimental study on three benchmark datasets demonstrates competent and/or higher classification performance (AID: 96.44%, NWPU: 94.32% and UCM: 99.04%) and robustness/stability (minimum standard deviation of 0.001) of the proposed approach.