<p>Current high-performance trackers typically follow an offline matching-based paradigm, where the target states in subsequent frames are inferred based on the target template from the initial frame, which makes it difficult for the tracker to adapt to target appearance changes and discriminate distractors on the fly. Moreover, they only learn to associate the tracked targets across temporal frames, but ignore to exploit the space-time correspondence of other contents. To alleviate above issues, we propose a novel tracking framework based on adjacent memory segmentation networks, termed as TAMS, which provides up-to-date target appearance information as well as the background contents around the target. Specifically, TAMS leverages the previous frame with the estimated target state to prompt the accuracy of current target estimation, enabling the tracker to adapt to the target appearance change well. Moreover, we introduce a self-supervised correspondence learning method to explore the intrinsic coherence in videos, which allows background pixels/patches in the current frame have the chance to find true correspondences in previous frames, reducing the cases of incorrect matching with the tracked target. Finally, in order to achieve an accurate description of the target, TAMS produces an accurate segmentation mask and bounding box jointly, that clearly distinguishing the target from background content. The extensive experimental results on seven mainstream visual tracking benchmarks show that the proposed tracker achieves promising tracking performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adjacent memory segmentation networks for robust visual tracking

  • Zhigang Shi,
  • Zhongyi Huang,
  • Zhiming Fang,
  • Feng Tang

摘要

Current high-performance trackers typically follow an offline matching-based paradigm, where the target states in subsequent frames are inferred based on the target template from the initial frame, which makes it difficult for the tracker to adapt to target appearance changes and discriminate distractors on the fly. Moreover, they only learn to associate the tracked targets across temporal frames, but ignore to exploit the space-time correspondence of other contents. To alleviate above issues, we propose a novel tracking framework based on adjacent memory segmentation networks, termed as TAMS, which provides up-to-date target appearance information as well as the background contents around the target. Specifically, TAMS leverages the previous frame with the estimated target state to prompt the accuracy of current target estimation, enabling the tracker to adapt to the target appearance change well. Moreover, we introduce a self-supervised correspondence learning method to explore the intrinsic coherence in videos, which allows background pixels/patches in the current frame have the chance to find true correspondences in previous frames, reducing the cases of incorrect matching with the tracked target. Finally, in order to achieve an accurate description of the target, TAMS produces an accurate segmentation mask and bounding box jointly, that clearly distinguishing the target from background content. The extensive experimental results on seven mainstream visual tracking benchmarks show that the proposed tracker achieves promising tracking performance.