<p>Existing methods, such as YOLOv8, face limitations in drone overhead views, including low accuracy for small objects, poor multi-scale adaptability, severe background interference, and high computational costs, which make them unsuitable for resource-constrained drones. To address these issues, this paper introduces the EBO-YOLO model for small object detection in such scenarios. Key improvements include: (1) Adding a small object detection layer to enhance semantic information and detection accuracy; (2) Replacing the original backbone with the EE network module (which fuses EfficientNetV1 and EMA attention) to reduce parameters and computational cost while maintaining sensitivity to small objects; (3) Integrating the BiFormer attention mechanism to improve multi-scale feature fusion and localization accuracy. Experiments on the VisDrone2019 Validation set show that the model uses 2.5<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4607_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> fewer parameters while achieving competitive performance on the TinyPerson and DOTAv1 datasets, especially in static scenes. However, its generalization to dynamic scenes and extremely tiny objects (below 10<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4607_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>10 pixels) requires further optimization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EBO-YOLO: small object detection from drone overhead views based on semantic refinement and detail acquisition

  • Xiaoyu Zhang,
  • Aziguli Wulamu,
  • Xi Guo,
  • Turdi Tohti

摘要

Existing methods, such as YOLOv8, face limitations in drone overhead views, including low accuracy for small objects, poor multi-scale adaptability, severe background interference, and high computational costs, which make them unsuitable for resource-constrained drones. To address these issues, this paper introduces the EBO-YOLO model for small object detection in such scenarios. Key improvements include: (1) Adding a small object detection layer to enhance semantic information and detection accuracy; (2) Replacing the original backbone with the EE network module (which fuses EfficientNetV1 and EMA attention) to reduce parameters and computational cost while maintaining sensitivity to small objects; (3) Integrating the BiFormer attention mechanism to improve multi-scale feature fusion and localization accuracy. Experiments on the VisDrone2019 Validation set show that the model uses 2.5 \(\times \) × fewer parameters while achieving competitive performance on the TinyPerson and DOTAv1 datasets, especially in static scenes. However, its generalization to dynamic scenes and extremely tiny objects (below 10 \(\times \) × 10 pixels) requires further optimization.