<p>The field of object detection has made significant advancements in detecting general objects. However, unmanned aerial vehicle (UAV) images commonly contain numerous small targets, and the amount of information they provide is limited. Therefore, the detection of UAV images remains a challenging task. To solve the issue of low detection accuracy in UAV images detection, a network called DMTNet is proposed in this paper. Specifically, it is proposed that Dual-Domain Adaptive Multi-scale Feature Fusion module can effectively address the issue of information loss during cross-layer fusion at the neck of the network and enhance context information by leveraging the joint action of spatial domain and frequency domain. By proposing Head with Hybrid Attention, more pixels are activated and redundant features are suppressed, enhancing the predictive potential of the detection head. Moreover, Non-stride DownSampling Convolution is designed to replace strided convolution in the feature extraction stage. This replacement helps preserve the fine-grained information of small targets, emphasizes the significant differences between feature image pixels, and enhances the model’s feature extraction capability. The results show that the proposed DMTNet demonstrates superior performance over other architectures within both the VisDrone2019, UAVDT and TT100K datasets. Compared to the YOLOv8s, the <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_3825_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="37" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{AP}}_{50}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mtext>AP</mtext> <mn>50</mn> </msub> </math></EquationSource> </InlineEquation> was improved by 8.9%, 2.4% and 6.4%, respectively. This approach notably improves the capability to detect small targets. All codes are available at <a href="https://github.com/sangxueting/DMTNet">https://github.com/sangxueting/DMTNet</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DMTNet: dual-domain adaptive multi-scale feature fusion network with transformer for small target detection

  • Yan Zhang,
  • Xueting Sang,
  • Yemei Sun,
  • Shudong Liu,
  • Shengpei Zhou

摘要

The field of object detection has made significant advancements in detecting general objects. However, unmanned aerial vehicle (UAV) images commonly contain numerous small targets, and the amount of information they provide is limited. Therefore, the detection of UAV images remains a challenging task. To solve the issue of low detection accuracy in UAV images detection, a network called DMTNet is proposed in this paper. Specifically, it is proposed that Dual-Domain Adaptive Multi-scale Feature Fusion module can effectively address the issue of information loss during cross-layer fusion at the neck of the network and enhance context information by leveraging the joint action of spatial domain and frequency domain. By proposing Head with Hybrid Attention, more pixels are activated and redundant features are suppressed, enhancing the predictive potential of the detection head. Moreover, Non-stride DownSampling Convolution is designed to replace strided convolution in the feature extraction stage. This replacement helps preserve the fine-grained information of small targets, emphasizes the significant differences between feature image pixels, and enhances the model’s feature extraction capability. The results show that the proposed DMTNet demonstrates superior performance over other architectures within both the VisDrone2019, UAVDT and TT100K datasets. Compared to the YOLOv8s, the \({\text{AP}}_{50}\) AP 50 was improved by 8.9%, 2.4% and 6.4%, respectively. This approach notably improves the capability to detect small targets. All codes are available at https://github.com/sangxueting/DMTNet.