<p>Real-time object detection is widely applied in fields such as autonomous driving and intelligent surveillance. However, it often faces a trade-off between computational efficiency and detection accuracy. Although Transformer-based end-to-end detection frameworks offer advantages in global modeling and optimization, their high computational costs and insufficient feature fusion limit both accuracy and real-time performance. To address these challenges, this paper proposes a novel object detection framework, PRepDETR, built upon the RT-DETR architecture. It introduces a lightweight partial convolution module to reduce redundant computations in the backbone. Additionally, a cross-scale feature fusion module, PRepGFPN, is developed by integrating partial convolution and reparameterization strategies, effectively enhancing inter-layer feature interaction. Furthermore, we present a novel bounding box regression loss function, Focaler-GIoU, which dynamically focuses on samples of varying difficulty, alleviates sample weighting imbalance, and improves the model’s generalization capability. Experimental results demonstrate that PRepDETR significantly enhances detection performance while maintaining real-time capability, achieving improvements of 2.4%/2.6% (mAP<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(_{\text {50}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mtext>50</mtext> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation>/mAP<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(_{\text {50-95}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mtext>50-95</mtext> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation>) on PASCAL VOC 2007, and 2.3%/2.5% on MS COCO 2017, with FPS increased by 8.0% and 9.8%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PRepDETR: a real-time detection transformer with partial convolutions and efficient cross-scale feature fusion

  • Bin Xiao,
  • Junchao Wen,
  • Xianjie He,
  • Yanxue Wu,
  • Xingpeng Zhang,
  • Shengtong Hu

摘要

Real-time object detection is widely applied in fields such as autonomous driving and intelligent surveillance. However, it often faces a trade-off between computational efficiency and detection accuracy. Although Transformer-based end-to-end detection frameworks offer advantages in global modeling and optimization, their high computational costs and insufficient feature fusion limit both accuracy and real-time performance. To address these challenges, this paper proposes a novel object detection framework, PRepDETR, built upon the RT-DETR architecture. It introduces a lightweight partial convolution module to reduce redundant computations in the backbone. Additionally, a cross-scale feature fusion module, PRepGFPN, is developed by integrating partial convolution and reparameterization strategies, effectively enhancing inter-layer feature interaction. Furthermore, we present a novel bounding box regression loss function, Focaler-GIoU, which dynamically focuses on samples of varying difficulty, alleviates sample weighting imbalance, and improves the model’s generalization capability. Experimental results demonstrate that PRepDETR significantly enhances detection performance while maintaining real-time capability, achieving improvements of 2.4%/2.6% (mAP \(_{\text {50}}\) 50 /mAP \(_{\text {50-95}}\) 50-95 ) on PASCAL VOC 2007, and 2.3%/2.5% on MS COCO 2017, with FPS increased by 8.0% and 9.8%, respectively.