Joint Frame and Event Object Tracking via Non-causal State Space Duality
摘要
RGB-Event object tracking has demonstrated significant potential in addressing challenging scenarios such as motion blur and high dynamic range, with its success largely hinging on the efficient fusion of visual information from both modalities. Existing Transformer-based methods achieve strong performance in multimodal learning but are hindered by high memory and computational costs, limiting practical deployment. Recently, State Space Model (SSM) has emerged as an effective mechanism for global token interactions with linear computational complexity. However, their inherent causal processing disrupts spatial structures in images, limiting their suitability for non-causal vision tasks. To address this issue, this paper proposes a novel RGB-Event tracking framework based on a non-causal format of State Space Duality (SSD). The framework comprises a backbone network for joint feature extraction and relation modeling, along with a fusion module to facilitate complementary integration of RGB and event modalities, achieving both efficiency and high performance. Extensive experiments on the FE108 and VisEvent datasets validate the effectiveness of the proposed method, showing significant improvements in tracking accuracy compared to state-of-the-art SSM-based trackers.