<p>Advances in underwater recording and processing systems have underscored the need for automated methods to detect and track small underwater objects accurately in imagery. However, the unique challenges posed by underwater optical images, including low contrast, color variations, and the presence of small objects, complicate detection and classification. To address these limitations, this research proposes a novel optimized deep-learning-based approach for underwater object detection and classification. In the preprocessing stage, UGIF-Net is used to restore natural color fidelity, mitigating distortions caused by light absorption and scattering. Contrast enhancement is then achieved through the perceptual dark channel prior (PDCP), improving the visibility and clarity of the images. Improved Conditional Generative adversarial Network (Improved CTGAN) is used to balance the data set. The LeViT approach is employed for robust feature extraction, capturing spatial and contextual information efficiently. For accurate detection and classification, a multi-branch fusion of spatial–temporal joint attention GCN (STJA–GCN) and DeepUWNet is applied, combining spatial–temporal dynamics with fine-grained segmentation. The chaotic-based tumbleweed optimization algorithm (CTOA) is utilized for hyperparameter optimization, ensuring optimal model performance. Finally, GradCAM++ is employed to visualize the critical image regions influencing the model's classification decisions, providing insights into its reasoning process. The novelty of this research lies in its comprehensive, integrated methodology that combines advanced color restoration, contrast enhancement, robust feature extraction, and hyperparameter optimization, alongside the use of GradCAM++ for model interpretability. The use of UGIF-Net, PDCP, and LeViT together, followed by a multi-branch fusion approach, presents a novel architecture that effectively handles underwater image challenges, improving both detection accuracy and the model's ability to provide transparent decision-making. Experimental results on four benchmark data sets highlight the superior performance of the proposed model, achieving mean average precision (mAP) scores of 98.42%, 98.43%, 98.78% and 99.01%. The proposed approach outperforms baseline YOLOv8 and other advanced non-YOLO methods by at least 5.2% and 5.3%, respectively. These findings confirm the effectiveness and robustness of the proposed network for underwater object detection in complex environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An optimized deep-learning framework for underwater object detection and classification using UGIF-Net and spatiotemporal attention

  • Jainabbi Banda,
  • J. Harikiran

摘要

Advances in underwater recording and processing systems have underscored the need for automated methods to detect and track small underwater objects accurately in imagery. However, the unique challenges posed by underwater optical images, including low contrast, color variations, and the presence of small objects, complicate detection and classification. To address these limitations, this research proposes a novel optimized deep-learning-based approach for underwater object detection and classification. In the preprocessing stage, UGIF-Net is used to restore natural color fidelity, mitigating distortions caused by light absorption and scattering. Contrast enhancement is then achieved through the perceptual dark channel prior (PDCP), improving the visibility and clarity of the images. Improved Conditional Generative adversarial Network (Improved CTGAN) is used to balance the data set. The LeViT approach is employed for robust feature extraction, capturing spatial and contextual information efficiently. For accurate detection and classification, a multi-branch fusion of spatial–temporal joint attention GCN (STJA–GCN) and DeepUWNet is applied, combining spatial–temporal dynamics with fine-grained segmentation. The chaotic-based tumbleweed optimization algorithm (CTOA) is utilized for hyperparameter optimization, ensuring optimal model performance. Finally, GradCAM++ is employed to visualize the critical image regions influencing the model's classification decisions, providing insights into its reasoning process. The novelty of this research lies in its comprehensive, integrated methodology that combines advanced color restoration, contrast enhancement, robust feature extraction, and hyperparameter optimization, alongside the use of GradCAM++ for model interpretability. The use of UGIF-Net, PDCP, and LeViT together, followed by a multi-branch fusion approach, presents a novel architecture that effectively handles underwater image challenges, improving both detection accuracy and the model's ability to provide transparent decision-making. Experimental results on four benchmark data sets highlight the superior performance of the proposed model, achieving mean average precision (mAP) scores of 98.42%, 98.43%, 98.78% and 99.01%. The proposed approach outperforms baseline YOLOv8 and other advanced non-YOLO methods by at least 5.2% and 5.3%, respectively. These findings confirm the effectiveness and robustness of the proposed network for underwater object detection in complex environments.