Modality balancing network for pedestrian detection based on cross-modal compensation fusion and multimodal feature alignment
摘要
Multimodal imaging technology improves the performance of pedestrian detection systems by complementing visible-light and infrared modalities, but problems such as modal imbalance and misalignment still need to be solved. Aiming at these issues, we propose an innovative modality balancing network for multimodal pedestrian detection to better integrate and collaborate the information of visible-light and infrared modalities. This network employs a cross-modal compensation fusion module to enhance image details through upsampling and downsampling operations and realize feature cross-utilization by sharing feature maps from different channels. It also utilizes a multimodal feature alignment mechanism to select complementary features according to lighting conditions, adaptively aligning the features of different modalities. Experimental results demonstrate that on the challenging KAIST multispectral pedestrian dataset, this network performs well in terms of detection accuracy and efficiency and is able to deal with the issue of lighting changes effectively.