Adaptive Text Feature Updating for Visual-Language Tracking
摘要
Visual-language tracking combines visual and textual information to improve the accuracy of object tracking in video sequences. However, existing methods utilize language descriptions only from the initial frame of the video sequence, leading to inaccuracies as the target’s appearance changes. To overcome this limitation, we propose a real-time updating method for language descriptions based on the target’s current features. Our approach incorporates a large visual-language model that continuously generates descriptions to maintain relevance and accuracy, as well as an update determination module to assess whether the current text’s quality requires refreshing. Additionally, we introduce a text fusion method to combine descriptions from different moments within the sequence, enhancing coherence and precision. Our method’s effectiveness is validated on OTB99, LaSOT, and TNL2K datasets, demonstrating superior performance and adaptability in various tracking scenarios.