Small object detection has widespread application in practical fields such as deep-sea archaeology and military early warning systems. However, achieving accurate and automated detection of small objects in complex scenes re-mains a significant challenge. Classical detection frameworks typically prioritize global image information and deprioritize small scale or extremely small-scale targets. Furthermore, continuous down-sampling in early net-work stages and feature fusion processes can result in loss of small object features, reducing the model’s ability to discern localized or extremely small scale targets. This limitation often leads to misdetections or false alarms in cluttered environments. To address these limitations, this paper proposes a scale enhanced small object detection framework based on the YOLOX architecture. Specifically, this paper introduced the triangular dilated convolution module in the baseline framework to mitigate the interference of redundant features during subsequent small object feature fusion calculations. At the feature fusion stage, this paper proposed a multi-scale hybrid attention mechanism and designed feature fusion architecture to efficiently capture global information for small objects with high efficiency. Comprehensive experiments demonstrate that our method achieves a 0.7% improvement in extremely small target detection. On the USD dataset, the proposed framework improves the detection rate of extremely small target by 1% and achieves 91.8% accuracy for standard scale small targets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TAM-YOLOX: Small Target Detection Based on Triangular Dilated Convolution Module and Multi-scale Hybrid Attention Mechanism

  • Jiaxing Wang,
  • Xuewei Li,
  • Yuquan Wu

摘要

Small object detection has widespread application in practical fields such as deep-sea archaeology and military early warning systems. However, achieving accurate and automated detection of small objects in complex scenes re-mains a significant challenge. Classical detection frameworks typically prioritize global image information and deprioritize small scale or extremely small-scale targets. Furthermore, continuous down-sampling in early net-work stages and feature fusion processes can result in loss of small object features, reducing the model’s ability to discern localized or extremely small scale targets. This limitation often leads to misdetections or false alarms in cluttered environments. To address these limitations, this paper proposes a scale enhanced small object detection framework based on the YOLOX architecture. Specifically, this paper introduced the triangular dilated convolution module in the baseline framework to mitigate the interference of redundant features during subsequent small object feature fusion calculations. At the feature fusion stage, this paper proposed a multi-scale hybrid attention mechanism and designed feature fusion architecture to efficiently capture global information for small objects with high efficiency. Comprehensive experiments demonstrate that our method achieves a 0.7% improvement in extremely small target detection. On the USD dataset, the proposed framework improves the detection rate of extremely small target by 1% and achieves 91.8% accuracy for standard scale small targets.