<p>Monocular depth estimation (MDE) plays a fundamental role in 3D scene modeling and related downstream applications. Existing weakly supervised methods for monocular depth estimation face two major challenges: (1) sparse supervision provides insufficient constraints for accurate depth learning; (2) the harmonization and fusion of global semantics with local details in the feature fusion paradigm need improvement to enhance depth boundaries and fine-grained accuracy. These limitations stem from passive learning frameworks and suboptimal feature aggregation strategies. To overcome these challenges, we introduce DepthRL, a novel weakly supervised method for monocular depth estimation, leveraging deep reinforcement learning. Unlike conventional methods, we reframe depth estimation as an optimization problem, utilizing reinforcement learning to determine the optimal policy for accurate depth prediction without the need for dense annotations. Another core of DepthRL lies in the Bi-directional Pyramid Policy Network (BPPN), where we effectively fuse global semantic information and fine-grained local details through multiscale feature interactions to generate high-quality depth maps. In this network, we introduce the Multiscale Cross Attention Feature Enhancement Module (MCFEM), which enhances the upward transfer of low-order features, allowing high-level semantics to be complemented by low-level local details. We also propose a novel feature fusion module, Symmetric Gated Attention Fusion (SGAFM), where the gated attention mechanism is used to suppress redundant and conflicting information, addressing semantic inconsistencies between different features. Experimental evaluations on the benchmark KITTI and NYU Depth v2 datasets, covering both indoor and outdoor environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DepthRL: a weakly supervised approach for monocular depth estimation using deep reinforcement learning

  • Han Chen,
  • Yongxiong Wang,
  • Jiayi Zhang,
  • Jiapeng Zhang,
  • Zhiqun Pan,
  • Shuwen Jia,
  • Shuai Huang

摘要

Monocular depth estimation (MDE) plays a fundamental role in 3D scene modeling and related downstream applications. Existing weakly supervised methods for monocular depth estimation face two major challenges: (1) sparse supervision provides insufficient constraints for accurate depth learning; (2) the harmonization and fusion of global semantics with local details in the feature fusion paradigm need improvement to enhance depth boundaries and fine-grained accuracy. These limitations stem from passive learning frameworks and suboptimal feature aggregation strategies. To overcome these challenges, we introduce DepthRL, a novel weakly supervised method for monocular depth estimation, leveraging deep reinforcement learning. Unlike conventional methods, we reframe depth estimation as an optimization problem, utilizing reinforcement learning to determine the optimal policy for accurate depth prediction without the need for dense annotations. Another core of DepthRL lies in the Bi-directional Pyramid Policy Network (BPPN), where we effectively fuse global semantic information and fine-grained local details through multiscale feature interactions to generate high-quality depth maps. In this network, we introduce the Multiscale Cross Attention Feature Enhancement Module (MCFEM), which enhances the upward transfer of low-order features, allowing high-level semantics to be complemented by low-level local details. We also propose a novel feature fusion module, Symmetric Gated Attention Fusion (SGAFM), where the gated attention mechanism is used to suppress redundant and conflicting information, addressing semantic inconsistencies between different features. Experimental evaluations on the benchmark KITTI and NYU Depth v2 datasets, covering both indoor and outdoor environments.