Cross-Modality Differential Feature Interaction for Multispectral Pedestrian Detection
摘要
Pedestrian detection is a vital element in computer vision, essential for applications like video surveillance and autonomous driving, requiring high levels of durability and adaptability from detection methods. However, conventional approaches often underperform in complex environments—characterized by low-light conditions and rapidly changing weather—thus limiting their practical utility. This study presents PAC-YOLO, a multimodal pedestrian detection approach based on the YOLOv8 architecture, which significantly enhances performance by strengthening effective feature interactions across RGB and infrared modalities. At the core of PAC-YOLO is the integration of a modality interaction-enhanced attention module and a differential feature fusion network. These components enable the dynamic recalibration of feature importance across RGB and IR inputs, thereby facilitating deep, efficient cross-modal feature integration. Additionally, by embedding Partial Convolution (PConv) into the YOLO backbone, PAC-YOLO further refines feature representations and improves detection accuracy. Extensive empirical evaluations on the KAIST, LLVIP, and FLIR benchmark datasets demonstrate that PAC-YOLO significantly advances accuracy, efficiency, and robustness, especially under low-light and occlusion conditions. This research thus provides an innovative, practically viable solution for multimodal pedestrian detection, offering substantial prospects for enhancing performance in surveillance and autonomous system applications.