Early Data Fusion in Low-Light Vision for YOLO
摘要
Object detection presents unique challenges in low-light conditions due to low contrast and limited feature visibility. This work examines the effectiveness of early data fusion for the methods of visible and infrared images within YOLO-based object detection. Several pressing strategies were used, including single-modal training, sequential fine-tuning, and the fused data approach. The YOLO11x model was chosen due to its state-of-the-art performance and adaptability. The experimental evaluation was performed based on the LLVIP database. The results suggest that early fusion, especially direct stacking, can match or cross the baseline model trained only on visual images. On the other hand, specific fusion methods such as Weighted Sum Fusion or the Sequential approach may degrade the performance. This highlights the importance of tailoring fusion strategies for this dataset and motivates further research on hybrid and late fusion techniques in multimodal object detection.