<p>Human detection in thermal imaging is a critical area of research for surveillance systems, especially under challenging conditions such as low visibility and nighttime. This study proposes a novel approach using a Block-recurrent Visual Transformer (BViT) integrated into a Feature Pyramid Network (FPN) for enhanced multi-scale feature extraction. By combining BViTFPN with the Task-aligned One-stage Object Detection (TOOD) head, the model achieves improved detection accuracy and computational efficiency by effectively capturing complex spatial and temporal data in thermal video. Evaluation on a challenging thermal video dataset demonstrates significant improvements in mean Average Precision (mAP), particularly for small and medium-sized objects. The proposed model achieves a mean mAP of 0.5222, with high precision at mAP@50 (0.9319) and mAP@75 (0.5317), while maintaining an efficient processing speed of 35.46 FPS. This research not only advances human detection in thermal imaging but also provides insights into leveraging transformer-based architecture for robust feature extraction and dynamic temporal analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Block-recurrent visual transformer for enhanced human detection in thermal imaging

  • Pham Cung Le Thien Vu,
  • Pham The Bao,
  • Tan Dat Trinh

摘要

Human detection in thermal imaging is a critical area of research for surveillance systems, especially under challenging conditions such as low visibility and nighttime. This study proposes a novel approach using a Block-recurrent Visual Transformer (BViT) integrated into a Feature Pyramid Network (FPN) for enhanced multi-scale feature extraction. By combining BViTFPN with the Task-aligned One-stage Object Detection (TOOD) head, the model achieves improved detection accuracy and computational efficiency by effectively capturing complex spatial and temporal data in thermal video. Evaluation on a challenging thermal video dataset demonstrates significant improvements in mean Average Precision (mAP), particularly for small and medium-sized objects. The proposed model achieves a mean mAP of 0.5222, with high precision at mAP@50 (0.9319) and mAP@75 (0.5317), while maintaining an efficient processing speed of 35.46 FPS. This research not only advances human detection in thermal imaging but also provides insights into leveraging transformer-based architecture for robust feature extraction and dynamic temporal analysis.