<p>The aim of exploiting deep learning methodologies for object detection is to localize and accurately classify pertinent object from video streams or sequences of images with minimal latency. Nevertheless, the identification and spatial positioning of non-salient objects, which occupy a small pixel proportion in images or are often occluded, remain challenging. In response to this challenge, we introduce an advanced methodology for object detection, hierarchical feature enhancement (HFE)-YOLO, which builds upon YOLOv10. Specifically, we first improve the feature fusion network of YOLOv10 by replacing the original path aggregation network (PANet) with a Hierarchical Semantic Fusion Network (HSFNet) featuring cross-layer asynchronous connections. This enhances the semantic information related to non-salient objects at the model output. Furthermore, we optimize the backbone network of the YOLOv10 baseline by introducing a responsive feature enhancement (RFE) module, which further strengthens the representation of features associated with non-salient objects. Finally, we refine the CIoU loss function used in YOLOv10 by proposing a novel Relative Intersection over Union (RIoU) loss, which ensures that the detection results more closely align with the actual ground truth. HFE-YOLO exhibits robust performance on two extensive public datasets, MS COCO 2017 and CrowdHuman. Relative to the baseline model YOLOv10, HFE-YOLO achieves an enhancement in mean average precision (mAP) by 2.8 and 2.5% on these datasets, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HFE-YOLO: a hierarchical feature enhancement approach for non-salient object detection

  • Chengang Dong,
  • Yongkang Ding,
  • Li Wang,
  • Jianwei Hu

摘要

The aim of exploiting deep learning methodologies for object detection is to localize and accurately classify pertinent object from video streams or sequences of images with minimal latency. Nevertheless, the identification and spatial positioning of non-salient objects, which occupy a small pixel proportion in images or are often occluded, remain challenging. In response to this challenge, we introduce an advanced methodology for object detection, hierarchical feature enhancement (HFE)-YOLO, which builds upon YOLOv10. Specifically, we first improve the feature fusion network of YOLOv10 by replacing the original path aggregation network (PANet) with a Hierarchical Semantic Fusion Network (HSFNet) featuring cross-layer asynchronous connections. This enhances the semantic information related to non-salient objects at the model output. Furthermore, we optimize the backbone network of the YOLOv10 baseline by introducing a responsive feature enhancement (RFE) module, which further strengthens the representation of features associated with non-salient objects. Finally, we refine the CIoU loss function used in YOLOv10 by proposing a novel Relative Intersection over Union (RIoU) loss, which ensures that the detection results more closely align with the actual ground truth. HFE-YOLO exhibits robust performance on two extensive public datasets, MS COCO 2017 and CrowdHuman. Relative to the baseline model YOLOv10, HFE-YOLO achieves an enhancement in mean average precision (mAP) by 2.8 and 2.5% on these datasets, respectively.