<p>Tiny object detection (TOD) is a pivotal yet challenging area in computer vision, marked by issues like limited pixel representation, extreme scale variations, occlusion, and noisy backgrounds. This survey provides a systematic and comprehensive review of TOD methodologies, tracing the transition from traditional convolutional neural network (CNN)-based models to state-of-the-art transformer-based architectures. Key limitations of CNNs, such as feature loss at smaller scales and sensitivity to cluttered environments, are addressed by attention mechanisms and multi-scale learning, foundational to transformer designs. Applications spanning autonomous driving, aerial surveillance, medical imaging, underwater detection, and security highlight the critical role of TOD in real-world scenarios. Detailed evaluations of models like YOLO, Faster R-CNN, DETR, Vision Transformers (ViTs), and hybrid frameworks across datasets such as MS COCO, TinyPersons, DeepLesion, DOTA, and UAV123 demonstrate notable advancements. Transformer-based models, including MS Transformer and HTDet, outperform their predecessors by capturing fine-grained features, enhancing robustness, and improving computational efficiency. This survey identifies pressing challenges, including dataset limitations like class imbalance, inadequate representation of tiny objects, and incomplete annotations, alongside evaluating TOD-specific metrics and dynamic object tracking. Future directions include the development of balanced and diverse datasets, weakly supervised learning frameworks, semi-supervised learning strategies, and real-time TOD optimization for edge computing. By integrating detailed comparative analyses and highlighting critical gaps, this survey offers a robust foundation for advancing TOD methodologies, driving enhanced detection accuracy, and expanding applicability in complex real-world environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive review of deep learning-based tiny object detection: challenges, strategies, and future directions

  • Muhamad Muzammul,
  • Xi Li

摘要

Tiny object detection (TOD) is a pivotal yet challenging area in computer vision, marked by issues like limited pixel representation, extreme scale variations, occlusion, and noisy backgrounds. This survey provides a systematic and comprehensive review of TOD methodologies, tracing the transition from traditional convolutional neural network (CNN)-based models to state-of-the-art transformer-based architectures. Key limitations of CNNs, such as feature loss at smaller scales and sensitivity to cluttered environments, are addressed by attention mechanisms and multi-scale learning, foundational to transformer designs. Applications spanning autonomous driving, aerial surveillance, medical imaging, underwater detection, and security highlight the critical role of TOD in real-world scenarios. Detailed evaluations of models like YOLO, Faster R-CNN, DETR, Vision Transformers (ViTs), and hybrid frameworks across datasets such as MS COCO, TinyPersons, DeepLesion, DOTA, and UAV123 demonstrate notable advancements. Transformer-based models, including MS Transformer and HTDet, outperform their predecessors by capturing fine-grained features, enhancing robustness, and improving computational efficiency. This survey identifies pressing challenges, including dataset limitations like class imbalance, inadequate representation of tiny objects, and incomplete annotations, alongside evaluating TOD-specific metrics and dynamic object tracking. Future directions include the development of balanced and diverse datasets, weakly supervised learning frameworks, semi-supervised learning strategies, and real-time TOD optimization for edge computing. By integrating detailed comparative analyses and highlighting critical gaps, this survey offers a robust foundation for advancing TOD methodologies, driving enhanced detection accuracy, and expanding applicability in complex real-world environments.