Temporal Context Aggregation for Visual Tracking Using Valid Frame and Knowledge Token
摘要
Temporal context information from consecutive frames is crucial for visual tracking tasks. Current popular trackers depend on the appearance information provided by the first frame template as the core feature without any further updates. In situations such as long-term tracking, the target’s appearance may differ significantly from the initial features, resulting in tracking failure. To tackle this problem, we propose TCATrack, a tracker aimed at ensuring effective online updates of the target’s appearance features throughout the tracking process. TCATrack aggregates temporal context in two ways: 1) using a valid frame for explicit appearance updates, and 2) using a temporal knowledge token for implicit appearance updates. The valid frame is updated online to provide the most recent and reliable appearance reference for tracking, while the temporal knowledge token iteratively transfers information about feature-rich areas from frame to frame. These two approaches to temporal context aggregation significantly improve the success rate and precision of tracking across various environments. Experimental results indicate that TCATrack surpasses existing methods in performance on multiple benchmark datasets.