Visual-language tracking combines visual and textual information to improve the accuracy of object tracking in video sequences. However, existing methods utilize language descriptions only from the initial frame of the video sequence, leading to inaccuracies as the target’s appearance changes. To overcome this limitation, we propose a real-time updating method for language descriptions based on the target’s current features. Our approach incorporates a large visual-language model that continuously generates descriptions to maintain relevance and accuracy, as well as an update determination module to assess whether the current text’s quality requires refreshing. Additionally, we introduce a text fusion method to combine descriptions from different moments within the sequence, enhancing coherence and precision. Our method’s effectiveness is validated on OTB99, LaSOT, and TNL2K datasets, demonstrating superior performance and adaptability in various tracking scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Text Feature Updating for Visual-Language Tracking

  • Xuexin Liu,
  • Zhuojun Zou,
  • Jie Hao

摘要

Visual-language tracking combines visual and textual information to improve the accuracy of object tracking in video sequences. However, existing methods utilize language descriptions only from the initial frame of the video sequence, leading to inaccuracies as the target’s appearance changes. To overcome this limitation, we propose a real-time updating method for language descriptions based on the target’s current features. Our approach incorporates a large visual-language model that continuously generates descriptions to maintain relevance and accuracy, as well as an update determination module to assess whether the current text’s quality requires refreshing. Additionally, we introduce a text fusion method to combine descriptions from different moments within the sequence, enhancing coherence and precision. Our method’s effectiveness is validated on OTB99, LaSOT, and TNL2K datasets, demonstrating superior performance and adaptability in various tracking scenarios.