Cross-Modality Fusion Deformable Transformer for Multispectral Object Detection
摘要
Different spectral images contain different information, and fusing multispectral information can enhance the performance considerably of object detection. We introduce a multimodal object detection network that can effectively achieve feature-level fusion of RGB images and infrared images with lightweight network. Firstly, we utilize a CrossModality Deformable Transformer (CDT) module to capture the latent relationship between the two modalities. Additionally, we design a dual-stream network structure to fuse light images and infrared images while balancing the model parameter volume. Experimental results showed that this method attains the top results on the VEDAI, FLIR, and LLVIP datasets.