In this paper, we introduce TGTrack, a novel visual tracking framework integrating Text Modality Autoregression and Generative Template Updating. TGTrack expands the latent feature space with an autoregressive decoder that models text modality information, encoding target trajectory coordinates and reconstructing update templates for temporal information. By converting trajectory coordinates into discrete tokens for a retention-based decoder, we enhance temporal modeling and awareness of the tracker. A novel generative template updating strategy is presented to handle challenges like object deformation and occlusion, reconstructing update templates directly instead of the previous discriminative approach. Experimental results demonstrate TGTrack’s competitiveness on benchmarks: achieving 72.2% \(SR_{0.75}\) on GOT-10k and 80.0% \(P_{Norm}\) on LaSOT, validating our framework’s effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TGTrack: Text Modality Autoregression and Generative Template Updating for Visual Object Tracking

  • Hao Xu,
  • Yiding Liang,
  • Haomiao Liu,
  • Chuhuai Yue,
  • Bo Ma

摘要

In this paper, we introduce TGTrack, a novel visual tracking framework integrating Text Modality Autoregression and Generative Template Updating. TGTrack expands the latent feature space with an autoregressive decoder that models text modality information, encoding target trajectory coordinates and reconstructing update templates for temporal information. By converting trajectory coordinates into discrete tokens for a retention-based decoder, we enhance temporal modeling and awareness of the tracker. A novel generative template updating strategy is presented to handle challenges like object deformation and occlusion, reconstructing update templates directly instead of the previous discriminative approach. Experimental results demonstrate TGTrack’s competitiveness on benchmarks: achieving 72.2% \(SR_{0.75}\) on GOT-10k and 80.0% \(P_{Norm}\) on LaSOT, validating our framework’s effectiveness.