<p>Addressing the critical challenges of small target detection in complex underwater environments, such as low visibility, feature redundancy, and multi-scale target variation. This paper proposes CDH-DETR, a lightweight real-time detection framework built on deep learning. The framework systematically incorporates a UGAN-based image enhancement module to restore color fidelity and improve clarity, while reconstructing the backbone network with ContextGuided Blocks to achieve efficient feature extraction alongside reduced computational complexity. Furthermore, it employs a dynamic upsampling (DySample) strategy, to preserve multi-level spatial details, and introduces a High-level Screening-feature Fusion Pyramid Network (HSFPN) for adaptive multi-scale feature integration. Experimental evaluations demonstrate that the proposed model attains a <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\textrm{mAP}\)</EquationSource> </InlineEquation>@0.5 of 90.0%, with a compact model size of 14.86 MB and a low computational cost of 44.2 GFLOPS. Deployed on a BlueROV2 platform with a Jetson Xavier NX, the system achieves an inference speed of 89.6 FPS, exhibiting robust performance under turbid and low-light underwater conditions. This study offers an efficient and scalable vision solution for underwater robotics, with significant potential in marine exploration and ecological monitoring applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning based lightweight real-time detection framework for small target in complex underwater environments

  • Zhe Dong,
  • Qing Yang,
  • HaoLin Chen,
  • Dexin Gao

摘要

Addressing the critical challenges of small target detection in complex underwater environments, such as low visibility, feature redundancy, and multi-scale target variation. This paper proposes CDH-DETR, a lightweight real-time detection framework built on deep learning. The framework systematically incorporates a UGAN-based image enhancement module to restore color fidelity and improve clarity, while reconstructing the backbone network with ContextGuided Blocks to achieve efficient feature extraction alongside reduced computational complexity. Furthermore, it employs a dynamic upsampling (DySample) strategy, to preserve multi-level spatial details, and introduces a High-level Screening-feature Fusion Pyramid Network (HSFPN) for adaptive multi-scale feature integration. Experimental evaluations demonstrate that the proposed model attains a \(\textrm{mAP}\) @0.5 of 90.0%, with a compact model size of 14.86 MB and a low computational cost of 44.2 GFLOPS. Deployed on a BlueROV2 platform with a Jetson Xavier NX, the system achieves an inference speed of 89.6 FPS, exhibiting robust performance under turbid and low-light underwater conditions. This study offers an efficient and scalable vision solution for underwater robotics, with significant potential in marine exploration and ecological monitoring applications.