<p>Pedestrian detection is a crucial task in computer vision, applicable in object tracking, video surveillance, and autonomous driving. Recent years have witnessed substantial advancements in pedestrian detection due to the fast evolution of deep learning in object detection. Nonetheless, obstacles such as inadequate detection accuracy persist, mostly because of varied pedestrian postures and intricate environments. This study proposes the RT-DETR-improved model to overcome these issues based on the real-time detection transformer (RT-DETR). First, we incorporate the high-low frequency (HiLo) attention into the encoder, therefore enhancing the model’s detection performance. Furthermore, we present a nonlinear feature fusion module that fuses information from various feature scales and contexts more successfully. We also introduce a novel loss function, InnerMPDIoU, to enhance detection efficacy in congested environments. To evaluate our model’s performance, extensive experiments are conducted on the CityPersons dataset. Compared to the baseline model, the RT-DETR-improved model attains a 4.2% enhancement in mAP50, a 2.0% improvement in mAP, a 2.2% rise in accuracy, and a 3.1% gain in recall. The results demonstrate that the proposed method exhibits superior detection accuracy and robustness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Improved Transformer-Based Model for Urban Pedestrian Detection

  • Tianyong Wu,
  • Xiang Li,
  • Qiuxuan Dong

摘要

Pedestrian detection is a crucial task in computer vision, applicable in object tracking, video surveillance, and autonomous driving. Recent years have witnessed substantial advancements in pedestrian detection due to the fast evolution of deep learning in object detection. Nonetheless, obstacles such as inadequate detection accuracy persist, mostly because of varied pedestrian postures and intricate environments. This study proposes the RT-DETR-improved model to overcome these issues based on the real-time detection transformer (RT-DETR). First, we incorporate the high-low frequency (HiLo) attention into the encoder, therefore enhancing the model’s detection performance. Furthermore, we present a nonlinear feature fusion module that fuses information from various feature scales and contexts more successfully. We also introduce a novel loss function, InnerMPDIoU, to enhance detection efficacy in congested environments. To evaluate our model’s performance, extensive experiments are conducted on the CityPersons dataset. Compared to the baseline model, the RT-DETR-improved model attains a 4.2% enhancement in mAP50, a 2.0% improvement in mAP, a 2.2% rise in accuracy, and a 3.1% gain in recall. The results demonstrate that the proposed method exhibits superior detection accuracy and robustness.