OVIF-YOLO: A Multispectral Object Detection Algorithm for Visible-Infrared Image Fusion
摘要
This study introduces OVIF-YOLO, an innovative dual-modal object detection framework developed upon YOLOv12. The proposed framework is purpose-built to achieve reliable multispectral detection in challenging environmental conditions. To tackle insufficient cross-modal fusion and performance degradation under adverse conditions (e.g., low light, blur), we introduce three core components: a Dynamic Feature Fusion (DFF) module for adaptive feature integration, a C3k2-SSM module with a visual state space mechanism to enhance long-range dependencies and small target detection, and an A2C2f-CAS module based on convolutional additive attention for efficient global context modeling. Extensive experiments on FLIR-Aligned, M3FD, and LLVIP benchmarks show state-of-the-art performance, achieving 89.8% mAP50 and 65.0% mAP50:95 on M3FD, outperforming existing methods in accuracy and robustness. The framework effectively leverages complementary visible and infrared information, offering a powerful solution for real-time multispectral detection.