<p>With the rapid technological progress, drones, or unmanned aerial vehicles (UAVs), have emerged among the most important artificial intelligence (AI)-powered systems. With their aerial perspective, mobility, and cost-effectiveness, they became crucial for advancing AI-driven visual perception in various sectors. However, implementing generic object detection algorithms on these resource-limited devices remains a complex challenge. Towards efficient and more UAV-adapted systems, this paper introduces the Faster Real-time Detector based on You Only Look Once (FRD-YOLO). FRD-YOLO presents different optimizations on the functional and architectural perspectives. To adapt the model to the vision context of UAVs, our key enhancements include the addition of a new layer for detecting tiny objects and the removal of the detection layer for large objects to emphasize the small and tiny targets. We also introduce a structure-aware integration of C3Ghost blocks, inspired by Ghost Convolutions and Cross Stage Partial Network-based layers, to reduce computational cost, and integrate the Convolutional Block Attention Module to enhance the recall. The FRD-YOLO model demonstrates superior detection performance with reduced size and computations across its five scaled versions. Notably, the evaluation on the challenging VisDroneDet2021 dataset reveals that the FRD-YOLO-x achieves 11.23% higher mean Average Precision (mAP50) than the baseline model, with 25.44% less computational cost, reaching up to 44 Frames Per Second (FPS). Additionally, the FRD-YOLO model showcases reliable embedded inference on the Jetson TX2, with FRD-YOLO-n achieving 25.18 FPS and FRD-YOLO-x reducing model size by 58.1%, confirming the architecture’s strength for UAV deployment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FRD-YOLO: a faster real-time object detector for aerial imagery

  • Ines Ben Rouighi,
  • Hajer Chtioui,
  • Imen Jegham,
  • Ihsen Alouani,
  • Anouar Ben Khalifa

摘要

With the rapid technological progress, drones, or unmanned aerial vehicles (UAVs), have emerged among the most important artificial intelligence (AI)-powered systems. With their aerial perspective, mobility, and cost-effectiveness, they became crucial for advancing AI-driven visual perception in various sectors. However, implementing generic object detection algorithms on these resource-limited devices remains a complex challenge. Towards efficient and more UAV-adapted systems, this paper introduces the Faster Real-time Detector based on You Only Look Once (FRD-YOLO). FRD-YOLO presents different optimizations on the functional and architectural perspectives. To adapt the model to the vision context of UAVs, our key enhancements include the addition of a new layer for detecting tiny objects and the removal of the detection layer for large objects to emphasize the small and tiny targets. We also introduce a structure-aware integration of C3Ghost blocks, inspired by Ghost Convolutions and Cross Stage Partial Network-based layers, to reduce computational cost, and integrate the Convolutional Block Attention Module to enhance the recall. The FRD-YOLO model demonstrates superior detection performance with reduced size and computations across its five scaled versions. Notably, the evaluation on the challenging VisDroneDet2021 dataset reveals that the FRD-YOLO-x achieves 11.23% higher mean Average Precision (mAP50) than the baseline model, with 25.44% less computational cost, reaching up to 44 Frames Per Second (FPS). Additionally, the FRD-YOLO model showcases reliable embedded inference on the Jetson TX2, with FRD-YOLO-n achieving 25.18 FPS and FRD-YOLO-x reducing model size by 58.1%, confirming the architecture’s strength for UAV deployment.