In conditions of low visibility, such as nighttime or restricted vision, the fusion of thermal and visible light features becomes particularly important. The Transformer technology has gained popularity in recent research for its ability to capture global information, showing better performance in feature fusion than traditional CNNs. However, the Transformer-based approaches often lead to high computational complexity during forward propagation. To solve this problem, we propose a Light Grouped Residual Transformer (LightGR-Transformer). It uses a single-layer network with channel-wise grouping and learnable residual connections, reducing the complexity of multi-layer Transformers. This design reduces computational cost while preserving important information during feature fusion. It prevents the loss of key features in deeper networks. Additionally, to improve the detection accuracy of small objects in low-light conditions, we introduce deformable convolution layers during the feature extraction stage. These layers dynamically adjust the receptive field of the convolution kernel, enhancing local detail capture. Our experiments on the FLIR dataset show that LightGR-Transformer improves mAP50 by 3.6% compared to existing state-of-the-art methods. On the KAIST dataset for nighttime detection, our network achieves the lowest \({MR}^{-2}\) score, reaching state-of-the-art performance. Our detector also reduces computational cost by 30%–40% compared to the most advanced models while maintaining top-level performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LightGR-Transformer: Light Grouped Residual Transformer for Multispectral Object Detection

  • Mingming Li,
  • Fei Wu,
  • Yinjie Wang

摘要

In conditions of low visibility, such as nighttime or restricted vision, the fusion of thermal and visible light features becomes particularly important. The Transformer technology has gained popularity in recent research for its ability to capture global information, showing better performance in feature fusion than traditional CNNs. However, the Transformer-based approaches often lead to high computational complexity during forward propagation. To solve this problem, we propose a Light Grouped Residual Transformer (LightGR-Transformer). It uses a single-layer network with channel-wise grouping and learnable residual connections, reducing the complexity of multi-layer Transformers. This design reduces computational cost while preserving important information during feature fusion. It prevents the loss of key features in deeper networks. Additionally, to improve the detection accuracy of small objects in low-light conditions, we introduce deformable convolution layers during the feature extraction stage. These layers dynamically adjust the receptive field of the convolution kernel, enhancing local detail capture. Our experiments on the FLIR dataset show that LightGR-Transformer improves mAP50 by 3.6% compared to existing state-of-the-art methods. On the KAIST dataset for nighttime detection, our network achieves the lowest \({MR}^{-2}\) score, reaching state-of-the-art performance. Our detector also reduces computational cost by 30%–40% compared to the most advanced models while maintaining top-level performance.