<p>Real-time object detection in continuously streaming image and video data demands that models remain both highly accurate and computationally frugal, yet existing lightweight detectors either sacrifice multi-scale robustness by narrowing network width or incur unnecessary overhead through scene-agnostic feature fusion strategies. To address these limitations, we propose Gated Adaptive Fusion Detector (GAF-Det), a single-stage detector whose neck replaces fixed-weight concatenation with a Lightweight Gated Feature Fusion (LGFF) module that dynamically computes per-channel fusion weights conditioned on both the semantic and spatial streams, and whose backbone employs a Depthwise Separable Downsampling (DSDown) block that combines a learnable depthwise separable branch with a max-pooling branch to preserve fine-grained spatial detail at roughly 88% fewer parameters than standard strided convolution in the depthwise separable branch. Integrated with a three-head PAFPN architecture and an InnerShape-IoU regression loss, GAF-Det achieves 56.1% AP<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(_{50}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mn>50</mn> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> on MS&#xa0;COCO with 3.1M parameters and 14.8&#xa0;GFLOPs, running at 36&#xa0;FPS on the NVIDIA Jetson AGX Xavier, demonstrating that input-adaptive feature fusion and efficient downsampling together yield a favorable accuracy–latency trade-off for general-purpose edge deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GAF-Det: gated adaptive fusion detector for lightweight real-time object detection

  • Xuxu Chu,
  • Wenyang Fan,
  • Yuanyuan Meng,
  • Qingyuan Zeng,
  • Juncen Guo

摘要

Real-time object detection in continuously streaming image and video data demands that models remain both highly accurate and computationally frugal, yet existing lightweight detectors either sacrifice multi-scale robustness by narrowing network width or incur unnecessary overhead through scene-agnostic feature fusion strategies. To address these limitations, we propose Gated Adaptive Fusion Detector (GAF-Det), a single-stage detector whose neck replaces fixed-weight concatenation with a Lightweight Gated Feature Fusion (LGFF) module that dynamically computes per-channel fusion weights conditioned on both the semantic and spatial streams, and whose backbone employs a Depthwise Separable Downsampling (DSDown) block that combines a learnable depthwise separable branch with a max-pooling branch to preserve fine-grained spatial detail at roughly 88% fewer parameters than standard strided convolution in the depthwise separable branch. Integrated with a three-head PAFPN architecture and an InnerShape-IoU regression loss, GAF-Det achieves 56.1% AP \(_{50}\) 50 on MS COCO with 3.1M parameters and 14.8 GFLOPs, running at 36 FPS on the NVIDIA Jetson AGX Xavier, demonstrating that input-adaptive feature fusion and efficient downsampling together yield a favorable accuracy–latency trade-off for general-purpose edge deployment.