A One-Step Transformer-based Adaptive Tracking Network for hyperspectral videos
摘要
Hyperspectral object tracking leverages cross-modality fusion between hyperspectral and pseudo-color data to enhance accuracy in complex scenarios. However, conventional trackers suffer from feature loss and high computational complexity due to their sequential three-step pipeline: extracting features, integrating target information per modality, and finally fusing cross-modality features. To address this, we propose the One-Step Transformer-based Adaptive Tracking Network (OSTAT-Net), which unifies feature extraction, target integration, and cross-modality fusion into a single step via concatenated self-attention. The core contributions of this work include: (1) A hierarchical hyperspectral pyramid embedding that converts hyperspectral cubes into token embeddings while preserving spectral information and mitigating overfitting; (2) A Domain-Adaptive Hyperspectral Adapter (DHA) with a bottleneck structure, enabling parameter-efficient knowledge transfer from RGB trackers to overcome data scarcity; (3) An Adaptive-Frequency Weighting Filter (AWF) that maps features to the frequency domain, decoupling and suppressing non-target noise in fused representations. Benchmark experiments demonstrate that OSTAT-Net achieves a high inference speed of 54.5 FPS and performs robustly in extensive comparisons, with its one-step design alleviating critical limitations of traditional multi-step frameworks.