Small object detection (SOD) remains a challenging problem in computer vision due to the small size and limited features of objects, making them difficult to distinguish from complex backgrounds. A key challenge is the ambiguity of structural features, often caused by the absence of object contours. To address this, we propose a novel approach, Adaptive Multi-scale Contextual Attention (AMCA), which effectively integrates region-level semantic context to enhance small blurred object detection. Unlike traditional methods that focus primarily on visual features or image-level semantic context, our approach exploits contextual information at the regional level to better capture object-environment relationships. The advantage of this approach is that it can not only exploit the object-environment-occurrence based contextual information, but also effectively reduce the disturbance caused by the complex and variable visual background features. This contextual fusion can significantly boost the model’s ability to discriminate small blurred objects from complex backgrounds. Experimental results on the DroneVehicle and Small COCO datasets show that our method outperforms existing CNN-based and DETR-like models, and one new state-of-the-art mAP scores (76.7% and 44.9%, respectively) can be achieved.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Small Blurred Object Detection by Region-Level Semantic Context Exploiting

  • Bingbin Zhou,
  • Yuhan Zhou,
  • Yan Luo,
  • Chao Shi,
  • Chongyang Zhang

摘要

Small object detection (SOD) remains a challenging problem in computer vision due to the small size and limited features of objects, making them difficult to distinguish from complex backgrounds. A key challenge is the ambiguity of structural features, often caused by the absence of object contours. To address this, we propose a novel approach, Adaptive Multi-scale Contextual Attention (AMCA), which effectively integrates region-level semantic context to enhance small blurred object detection. Unlike traditional methods that focus primarily on visual features or image-level semantic context, our approach exploits contextual information at the regional level to better capture object-environment relationships. The advantage of this approach is that it can not only exploit the object-environment-occurrence based contextual information, but also effectively reduce the disturbance caused by the complex and variable visual background features. This contextual fusion can significantly boost the model’s ability to discriminate small blurred objects from complex backgrounds. Experimental results on the DroneVehicle and Small COCO datasets show that our method outperforms existing CNN-based and DETR-like models, and one new state-of-the-art mAP scores (76.7% and 44.9%, respectively) can be achieved.