<p>(UAVs) play a critical role in traffic flow monitoring and rapid accident response. However, this task remains challenging due to variable high-altitude viewpoints, complex environmental interference, and limitations in algorithmic efficiency. To address these issues, a lightweight vehicle detection model is developed based on UAV aerial imagery. In the backbone, a Receptive-Field Attention Convolution (RFAConv) module is introduced to retain detailed features during the downsampling process. To handle complex environments such as illumination changes and dense occlusion, a Cross-Stage Partial Fusion Bi-Level Routing Attention Network (CSP-BLRAN) module is employed to enhance contextual feature representation and suppress false and missed detections. In the neck, a multi-layer selective feature fusion pyramid (MS-FPN) is constructed to perform attention-based filtering on high-level semantic features and low-level spatial details, followed by feature enhancement via multiplication and global semantic refinement through residual connections. For the detection head, a Generalized Wasserstein Distance Loss (GWDLoss) function is proposed to quantify positional and scale discrepancies, improving the adaptability of bounding box regression to geometric variations. Extensive experiments on the VisDrone and Vehicle datasets demonstrate that the proposed method reduces the number of parameters by 22.9% and achieves mAP@50–95 improvements of 4.5% and 7.3%, respectively, over YOLO11n. The approach also surpasses mainstream models such as YOLO11s with significantly fewer parameters, confirming its effectiveness and practical value in resource-constrained low-altitude UAV-based detection tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vehicle detection method based on multi-layer selective feature for UAV aerial images

  • Yinbao Ma,
  • Yuyu Meng,
  • Jiuyuan Huo

摘要

(UAVs) play a critical role in traffic flow monitoring and rapid accident response. However, this task remains challenging due to variable high-altitude viewpoints, complex environmental interference, and limitations in algorithmic efficiency. To address these issues, a lightweight vehicle detection model is developed based on UAV aerial imagery. In the backbone, a Receptive-Field Attention Convolution (RFAConv) module is introduced to retain detailed features during the downsampling process. To handle complex environments such as illumination changes and dense occlusion, a Cross-Stage Partial Fusion Bi-Level Routing Attention Network (CSP-BLRAN) module is employed to enhance contextual feature representation and suppress false and missed detections. In the neck, a multi-layer selective feature fusion pyramid (MS-FPN) is constructed to perform attention-based filtering on high-level semantic features and low-level spatial details, followed by feature enhancement via multiplication and global semantic refinement through residual connections. For the detection head, a Generalized Wasserstein Distance Loss (GWDLoss) function is proposed to quantify positional and scale discrepancies, improving the adaptability of bounding box regression to geometric variations. Extensive experiments on the VisDrone and Vehicle datasets demonstrate that the proposed method reduces the number of parameters by 22.9% and achieves mAP@50–95 improvements of 4.5% and 7.3%, respectively, over YOLO11n. The approach also surpasses mainstream models such as YOLO11s with significantly fewer parameters, confirming its effectiveness and practical value in resource-constrained low-altitude UAV-based detection tasks.