Block-recurrent visual transformer for enhanced human detection in thermal imaging
摘要
Human detection in thermal imaging is a critical area of research for surveillance systems, especially under challenging conditions such as low visibility and nighttime. This study proposes a novel approach using a Block-recurrent Visual Transformer (BViT) integrated into a Feature Pyramid Network (FPN) for enhanced multi-scale feature extraction. By combining BViTFPN with the Task-aligned One-stage Object Detection (TOOD) head, the model achieves improved detection accuracy and computational efficiency by effectively capturing complex spatial and temporal data in thermal video. Evaluation on a challenging thermal video dataset demonstrates significant improvements in mean Average Precision (mAP), particularly for small and medium-sized objects. The proposed model achieves a mean mAP of 0.5222, with high precision at mAP@50 (0.9319) and mAP@75 (0.5317), while maintaining an efficient processing speed of 35.46 FPS. This research not only advances human detection in thermal imaging but also provides insights into leveraging transformer-based architecture for robust feature extraction and dynamic temporal analysis.