RGB-T Tracking via Region Filtering-Fusion and Prompt Learning
摘要
In UAV tracking systems, both infrared (IR) and visible light data contain significant redundant information. Furthermore, data quality fluctuates between these modalities. Low-quality features can degrade the fusion features and subsequently affect the reliability of tracking results. To address this issue, we propose a background removal network based on target confidence, aiming to eliminate redundant and low-quality information from IR and visible light images. Firstly, information from different modalities interacts within an encoder. The features generated through this interaction undergo scoring and selection processes, where regions with higher scores are retained, and regions with lower scores, indicative of low-quality redundancy, are removed. During target tracking, previous target position information strongly correlates with the current frame’s position. Thus, we introduce tokenized previous frame’s target position information as network prompts to guide attention to the previous target area. Combining historical target position information and visual encoder outputs, the decoder progressively decodes the current position information. Our algorithm based on visual and positional cues achieves superior performance and faster speeds compared to traditional visual-only tracking algorithms across several RGBT datasets, and has been successfully deployed in UAV systems.