<p>In recent years, underwater image object detection has become a hot research topic in computer vision. Due to the unique underwater environment, conventional object detection models tend to perform poorly. Specifically, the attenuation of light in water leads to color distortion and reduced contrast. Suspended particles in water cause scattering, resulting in image blurring and further color issues. Water absorbs different wavelengths of light, with red being the first to disappear, making images predominantly blue or green and severely disrupting color channels. Additionally, the lack of underwater datasets hinders the performance of object detection models in real-world applications, with most datasets being generated by adding noise to existing image datasets for training purposes. Current approaches primarily focus on independent preprocessing of underwater images to enhance their quality. The commonly used models are predominantly two-stage architectures based on RCNN, which improve accuracy at the cost of reduced frame rates. This paper proposes an underwater object detection network based on a single-stage YOLO model. By introducing a coordinate attention mechanism and an improved module based on transformer blocks, the proposed model reduces computational cost while increasing detection frame rates. Experimental results demonstrate that the model achieves both higher accuracy and computational efficiency, with <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1785_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="50" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{mAP}}_{{50}}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mtext>mAP</mtext> <mn>50</mn> </msub> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1785_Article_IEq2.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="71" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{mAP}}_{{50}-{95}}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mtext>mAP</mtext> <mrow> <mn>50</mn> <mo>-</mo> <mn>95</mn> </mrow> </msub> </math></EquationSource> </InlineEquation> reaching 96.8% and 81.2%, respectively. Moreover, the model’s GFlops is only 24.6, representing a 14% improvement over the baseline model. Our proposed model achieves superior computational efficiency, demonstrating significantly faster inference speeds and lower computational requirements compared to state-of-the-art approaches, while maintaining competitive detection accuracy on benchmark datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FA-YOLO: a YOLO model based on attention mechanism and feature fusion for object detection in multi-classification underwater datasets

  • Haiyong Wang,
  • Yi Zhou

摘要

In recent years, underwater image object detection has become a hot research topic in computer vision. Due to the unique underwater environment, conventional object detection models tend to perform poorly. Specifically, the attenuation of light in water leads to color distortion and reduced contrast. Suspended particles in water cause scattering, resulting in image blurring and further color issues. Water absorbs different wavelengths of light, with red being the first to disappear, making images predominantly blue or green and severely disrupting color channels. Additionally, the lack of underwater datasets hinders the performance of object detection models in real-world applications, with most datasets being generated by adding noise to existing image datasets for training purposes. Current approaches primarily focus on independent preprocessing of underwater images to enhance their quality. The commonly used models are predominantly two-stage architectures based on RCNN, which improve accuracy at the cost of reduced frame rates. This paper proposes an underwater object detection network based on a single-stage YOLO model. By introducing a coordinate attention mechanism and an improved module based on transformer blocks, the proposed model reduces computational cost while increasing detection frame rates. Experimental results demonstrate that the model achieves both higher accuracy and computational efficiency, with \({\text{mAP}}_{{50}}\) mAP 50 and \({\text{mAP}}_{{50}-{95}}\) mAP 50 - 95 reaching 96.8% and 81.2%, respectively. Moreover, the model’s GFlops is only 24.6, representing a 14% improvement over the baseline model. Our proposed model achieves superior computational efficiency, demonstrating significantly faster inference speeds and lower computational requirements compared to state-of-the-art approaches, while maintaining competitive detection accuracy on benchmark datasets.