<p>Constructing high-quality contextual information from video sequences is of critical importance for visual object tracking. However, existing methods typically use only a few markers to directly convey this information. This might introduce irrelevant information, thereby misleading the tracking of subsequent frames. To effectively address this issue, we propose a new video-level visual object tracking framework called Dynamic Adaptive Sparse Mamba for Video-Level Tracking (DASM). Its core component is the Attention-Guided Contextual Information Fusion Module (AGCIF). We further optimize the Mamba-based contextual propagation mechanism. Mamba is unable to independently evaluate the quality of new information when integrating contextual information. Therefore, we embedded a sparse attention mechanism in the hidden layer of Mamba, so that all new context information could only be fused after passing through the sparse attention mechanism. We utilize sparse attention to enable the Mamba layer to focus on local features strongly related to the target, enhancing the Mamba’s ability to filter out key information from mixed features. While retaining spatial neighborhood associations, the proposed design enhances the model’s spatiotemporal modeling ability for the target and promotes the utilization and downstream propagation of high-quality contextual information. However, in most existing Mamba-based trackers, the current-frame mixed responses are directly written into the memory update process, so target-relevant information and background-dominated responses may be accumulated together. Therefore, the remaining challenge is not only how to propagate more context, but also how to propagate higher-quality context in a controlled manner. This observation motivates our DASM framework, which introduces target-aware sparse filtering before state updating and thus transforms contextual propagation from passive accumulation into quality-controlled memory propagation. Furthermore, we have also proposed a Dynamic Video Segment Scheduling Module (DVSS), This module uses a multi-layer perceptron to evaluate recent multi-frame tracking results, as the basis for making decisions regarding the length of the input continuous video clips. This design ensures that the tracker captures sufficient high-quality information while balancing tracking accuracy and speed. Our tracker has undergone extensive experimentation across multiple datasets, validating its outstanding performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic adaptive sparse Mamba for video-level tracking

  • Huanlong Zhang,
  • Boyan Li,
  • Yao Yu Chen,
  • Jie Zhang,
  • Liao Zhu,
  • Junlong Gao,
  • Dayong Xu

摘要

Constructing high-quality contextual information from video sequences is of critical importance for visual object tracking. However, existing methods typically use only a few markers to directly convey this information. This might introduce irrelevant information, thereby misleading the tracking of subsequent frames. To effectively address this issue, we propose a new video-level visual object tracking framework called Dynamic Adaptive Sparse Mamba for Video-Level Tracking (DASM). Its core component is the Attention-Guided Contextual Information Fusion Module (AGCIF). We further optimize the Mamba-based contextual propagation mechanism. Mamba is unable to independently evaluate the quality of new information when integrating contextual information. Therefore, we embedded a sparse attention mechanism in the hidden layer of Mamba, so that all new context information could only be fused after passing through the sparse attention mechanism. We utilize sparse attention to enable the Mamba layer to focus on local features strongly related to the target, enhancing the Mamba’s ability to filter out key information from mixed features. While retaining spatial neighborhood associations, the proposed design enhances the model’s spatiotemporal modeling ability for the target and promotes the utilization and downstream propagation of high-quality contextual information. However, in most existing Mamba-based trackers, the current-frame mixed responses are directly written into the memory update process, so target-relevant information and background-dominated responses may be accumulated together. Therefore, the remaining challenge is not only how to propagate more context, but also how to propagate higher-quality context in a controlled manner. This observation motivates our DASM framework, which introduces target-aware sparse filtering before state updating and thus transforms contextual propagation from passive accumulation into quality-controlled memory propagation. Furthermore, we have also proposed a Dynamic Video Segment Scheduling Module (DVSS), This module uses a multi-layer perceptron to evaluate recent multi-frame tracking results, as the basis for making decisions regarding the length of the input continuous video clips. This design ensures that the tracker captures sufficient high-quality information while balancing tracking accuracy and speed. Our tracker has undergone extensive experimentation across multiple datasets, validating its outstanding performance.