A Review of Infrared and Visible Light Image Fusion Algorithms
摘要
With the advancement of autonomous driving technology, the integration of infrared and visible light imagery has emerged as a critical component for advanced driver-assistance systems (ADAS) and autonomous driving applications. Infrared imagery facilitates robust target detection under low-light conditions and adverse weather environments, whereas visible light imagery offers detailed textural information. The integration of these two categories of images significantly enhances environmental perception in terms of precision and robustness. This paper systematically reviews recent methodologies for infrared and visible light image fusion within the context of driving assistance, emphasizing deep learning-based approaches such as autoencoders, convolutional neural networks, generative adversarial networks, Transformer architectures, and emerging multimodal foundation models. In contrast to conventional image fusion methods grounded in mathematical models, deep learning-based approaches demonstrate superior performance in managing complex scenarios involving multiple targets, owing to their capacity for automatic feature extraction and fusion strategy optimization. The paper further assesses fusion metrics, presents widely-used datasets (e.g., LLVIP, RoadScene, and M3FD), underscores the constraints of existing methodologies, and proposes future research trajectories for multimodal and foundation-model-driven fusion paradigms.