Non-aligned RGBT Tracking via Joint Temporal-Iterated Homography Estimation and Multimodal Transformer Fusion
摘要
Existing RGBT tracking methods have limitations in practical applications due to their reliance on spatial-aligned videos, which often require either elaborate platform design or high-cost manual alignment. To address this issue, we propose a Non-Aligned RGBT Tracker (NAT) that can effectively utilizes both weakly-aligned and non-aligned data, enabling it to be trained and tested on both types of data. Our method consists of two key components, including the temporal-iterated homography estimation module and the multimodal transformer fusion module. The temporal-iterated homography estimation module learns a transformation using temporal knowledge. By considering the continuity of homography changes in multimodal video sequences, this module uses an iterated prediction method that leverages the guidance of predicted transformation parameters from previous frames. This enables stable, accurate, and robust homography estimation in weakly-aligned and non-aligned scenarios without pre-alignment, making it practical for applications. The multimodal transformer fusion module aims to capture complementary information from multiple modalities by exploiting the powerful global modeling capability of the transformer. The entire framework can be trained end-to-end and is evaluated on both weakly-aligned and non-aligned RGBT datasets, and the results suggest our NAT outperforms state-of-the-art methods on five RGBT tracking benchmarks. Our approach opens up the practical applications of RGBT tracking research.