Efficient Object Detection in Fused Visual and Infrared Spectra for Edge Platforms
摘要
Image fusion is an important task in computer vision, aiming to combine information from multiple modalities to enhance overall perception and understanding. This paper presents an innovative approach to image fusion, specifically focusing on the fusion of RGB and thermal image embeddings to enhance object recognition and detection performance. To achieve this, we introduce two key components, FusionConv and SepThermalConv layers to the YOLO object detection network, as well as modified FusionC3 layer. The FusionConv layer effectively integrates RGB and thermal image features by leveraging multimodal embeddings with use of fusion parameter. Similarly, the SepThermalConv layer optimizes the processing of thermal information by incorporating separate branches for enhanced representation. Extensive experiments conducted on a custom dataset demonstrate significant performance gains achieved by our fusion method compared to using individual modalities in isolation. Our results highlight the potential of multimodal fusion techniques to improve object detection and perception of complex scenes by effectively combining RGB and thermal images.