Self-supervised monocular depth estimation is a critical task in computer vision. Existing methods can be typically categorized into discriminative-based and generative-based methods according to different data modeling approaches. Discriminative-based methods are distinguished by high accuracy, while generative-based methods are notable for superior robustness. Given that images captured in real-world scenarios are inevitably influenced by various external factors, it is essential to develop a robust and accurate depth estimation algorithm. However, there is limited research on the balance of robustness and accuracy by exploring the interactions between discriminative and generative networks. We propose a generative diffusion-based self-supervised monocular depth estimation algorithm guided by discriminative networks and incorporate a depth interaction constraint. We utilize discriminative networks to optimize image-guided information for the denoising process within the diffusion model. This approach seamlessly combines great robustness with high accuracy. Additionally, to reduce the impact of low texture regions on the reprojection photometric loss, we design a texture-aware discriminatory mask module. This module strengthens the constraint capability of the photometric consistency. We conduct experiments on the KITTI and Make3D datasets. The results demonstrate that our method successfully balances accuracy and robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Discriminative-Guided Diffusion-Based Self-supervised Monocular Depth Estimation

  • Runze Liu,
  • Guanghui Zhang,
  • Dongchen Zhu,
  • Lei Wang,
  • Xiaolin Zhang,
  • Jiamao Li

摘要

Self-supervised monocular depth estimation is a critical task in computer vision. Existing methods can be typically categorized into discriminative-based and generative-based methods according to different data modeling approaches. Discriminative-based methods are distinguished by high accuracy, while generative-based methods are notable for superior robustness. Given that images captured in real-world scenarios are inevitably influenced by various external factors, it is essential to develop a robust and accurate depth estimation algorithm. However, there is limited research on the balance of robustness and accuracy by exploring the interactions between discriminative and generative networks. We propose a generative diffusion-based self-supervised monocular depth estimation algorithm guided by discriminative networks and incorporate a depth interaction constraint. We utilize discriminative networks to optimize image-guided information for the denoising process within the diffusion model. This approach seamlessly combines great robustness with high accuracy. Additionally, to reduce the impact of low texture regions on the reprojection photometric loss, we design a texture-aware discriminatory mask module. This module strengthens the constraint capability of the photometric consistency. We conduct experiments on the KITTI and Make3D datasets. The results demonstrate that our method successfully balances accuracy and robustness.