<p>The depth information extracted from underwater images plays a crucial role in computer vision tasks such as underwater target tracking, image modeling and analysis, and precise target localization. Unlike terrestrial depth estimation, underwater monocular depth estimation faces unique challenges: low contrast and blurring effects impair the model’s ability to extract edge details, while uneven illumination causes inconsistent luminance distribution, making it difficult to establish stable depth mapping relationships. To address these challenges, this paper proposes an underwater monocular depth estimation model based on local perception transformer and global–local context fusion. Specifically, the local perception unit extracts fine-grained textures and local spatial features, while the vision Transformer component requires GPU acceleration to efficiently capture long-range dependencies and compensate for semantic information loss caused by local blurring. The global–local self-attention mechanism further demands high-performance computing resources to adaptively adjust regional attention weights in large-scale underwater datasets, thereby mitigating the impact of luminance variations on structural understanding and improving depth estimation accuracy. Experiments on the publicly available FLSea underwater dataset demonstrate that the model achieves an AbsRel value of 0.147, representing improvements of 30%, 55.6%, and 81.2% over AdaBins, PixelFormer, and LapDepth, respectively. This efficient yet computationally demanding architecture also exhibits excellent generalization performance on the Sea-thru dataset while maintaining practical deployability and meeting real-time processing requirements through high-performance computing capabilities.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LPTMono: monocular depth estimation for underwater images using local perception transformer and global–local context fusion

  • Dingshuo Liu,
  • Yiran Liu,
  • Beibei Li,
  • Qingling Duan

摘要

The depth information extracted from underwater images plays a crucial role in computer vision tasks such as underwater target tracking, image modeling and analysis, and precise target localization. Unlike terrestrial depth estimation, underwater monocular depth estimation faces unique challenges: low contrast and blurring effects impair the model’s ability to extract edge details, while uneven illumination causes inconsistent luminance distribution, making it difficult to establish stable depth mapping relationships. To address these challenges, this paper proposes an underwater monocular depth estimation model based on local perception transformer and global–local context fusion. Specifically, the local perception unit extracts fine-grained textures and local spatial features, while the vision Transformer component requires GPU acceleration to efficiently capture long-range dependencies and compensate for semantic information loss caused by local blurring. The global–local self-attention mechanism further demands high-performance computing resources to adaptively adjust regional attention weights in large-scale underwater datasets, thereby mitigating the impact of luminance variations on structural understanding and improving depth estimation accuracy. Experiments on the publicly available FLSea underwater dataset demonstrate that the model achieves an AbsRel value of 0.147, representing improvements of 30%, 55.6%, and 81.2% over AdaBins, PixelFormer, and LapDepth, respectively. This efficient yet computationally demanding architecture also exhibits excellent generalization performance on the Sea-thru dataset while maintaining practical deployability and meeting real-time processing requirements through high-performance computing capabilities.