Learning Explicit Modulation Vectors for Disentangled Transformer Attention-Based RGB-D Visual Tracking
摘要
Effective fusion of RGB-D multi-modal features is crucial for RGB-D object tracking. Existing fusion methods mainly guide the interaction of RGB and depth by dense attention, but such formulation relies only on independent spatial token attributes, without considering the correspondences among channel slices. To address this limitation, we propose a spatial and channel disentangled attention mechanism, providing dual guidance on spatial and channel relevance for RGB-D fusion. Simultaneously, to deal with potential erroneous attention, we exploit the explicit modulation vectors to weaken less relevant spatial and channel features. Drawing on this, we design an adaptive architecture by weakening the significance of less confident intra-modal features and amplifying the supportive cross-modal features. The experimental results on four standard RGB-D benchmarking datasets, i.e., ARKitTrack, DepthTrack, RGBD1K, and CDTB, confirm the merit of our approach in adaptive fusion, outperforming existing state-of-the-art solutions.