Effective fusion of RGB-D multi-modal features is crucial for RGB-D object tracking. Existing fusion methods mainly guide the interaction of RGB and depth by dense attention, but such formulation relies only on independent spatial token attributes, without considering the correspondences among channel slices. To address this limitation, we propose a spatial and channel disentangled attention mechanism, providing dual guidance on spatial and channel relevance for RGB-D fusion. Simultaneously, to deal with potential erroneous attention, we exploit the explicit modulation vectors to weaken less relevant spatial and channel features. Drawing on this, we design an adaptive architecture by weakening the significance of less confident intra-modal features and amplifying the supportive cross-modal features. The experimental results on four standard RGB-D benchmarking datasets, i.e., ARKitTrack, DepthTrack, RGBD1K, and CDTB, confirm the merit of our approach in adaptive fusion, outperforming existing state-of-the-art solutions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Explicit Modulation Vectors for Disentangled Transformer Attention-Based RGB-D Visual Tracking

  • Yifan Pan,
  • Tianyang Xu,
  • Xue-Feng Zhu,
  • Xiaoqing Luo,
  • Xiao-Jun Wu,
  • Josef Kittler

摘要

Effective fusion of RGB-D multi-modal features is crucial for RGB-D object tracking. Existing fusion methods mainly guide the interaction of RGB and depth by dense attention, but such formulation relies only on independent spatial token attributes, without considering the correspondences among channel slices. To address this limitation, we propose a spatial and channel disentangled attention mechanism, providing dual guidance on spatial and channel relevance for RGB-D fusion. Simultaneously, to deal with potential erroneous attention, we exploit the explicit modulation vectors to weaken less relevant spatial and channel features. Drawing on this, we design an adaptive architecture by weakening the significance of less confident intra-modal features and amplifying the supportive cross-modal features. The experimental results on four standard RGB-D benchmarking datasets, i.e., ARKitTrack, DepthTrack, RGBD1K, and CDTB, confirm the merit of our approach in adaptive fusion, outperforming existing state-of-the-art solutions.