VAFTrack: asynchronous feature fusion via visual receptive weighted key-value perceptual for visual tracking
摘要
To address the challenges of inadequate fusion and redundancy in the interaction between template and search features within visual encoders, we propose an Asynchronous Fusion Tracking Model based on Visual Receptive Weighted Key-Value Perceptual(VAFTrack). Initially, we introduce a Visual Receptive Weighted Key-Value Perceptual Fusion Module(VPFM) that enhances the matching accuracy between template and search features by sequentially integrating Contextual Spatial Perception Module(CSPM) and Contextual Channel Perception Module(CCPM). This approach improves the model’s target awareness and adaptability to variations in scenes. Subsequently, we design a Spatial Optimization Enhancement Module(SOEM) and a Channel Optimization Enhancement Module(COEM) to mitigate redundancy and augment information density by optimizing spatial and channel representations within feature maps. These enhancement bolsters the model’s robustness against occlusion and complex backgrounds, thereby strengthening its discriminative capacity for target tracking. The proposed model achieves state-of-the-art performance, with accuracies of 81.9% on TrackingNet and 73.6% on LaSOT, and an overlap rate of 73.6% on the GOT-10k dataset.