DynaMix-VLT: Dynamic Multi-head Attention and Semantic-guided Modal Mixing for Robust Vision-language Tracking
摘要
As a fundamental vision task, tracking gains fewer benefits from recent flourishing vision-language (VL) learning. Most existing VL trackers treat language as auxiliary information for vision, ignoring the interaction learning and complementarity between them. In this paper, a novel VL tracking framework named DynaMix vision-language tracker (DVLT) is developed to learn joint feature extraction and interaction based on the transformer backbone. First, a dynamic multi-head attention (DMHA) mechanism is designed to alleviate the low-rank bottleneck of traditional multi-head attention by employing a collaboration function among heads at the level of the attention matrix. Then, a trainable shared composition map is proposed for dynamically synthesizing new attention heads to maintain computational speed and simultaneously improve the attention matrix rank. In addition, to further enhance the ability of model representation, a modal mixing strategy is utilized to make the extracted features highly target-aware by assigning weights to the visual tokens under language guidance. Finally, a semantic location head is introduced that utilizes reference information to reduce visual ambiguity and mines target features by introducing high-level semantic information based on a similarity function. Extensive experiments are conducted on multiple language-assisted tracking datasets, including OTB99-LANG, TNL2K, LaSOT and LaSOText, to demonstrate the superiority of the proposed tracker compared with existing VL trackers.