Discriminative-Guided Diffusion-Based Self-supervised Monocular Depth Estimation
摘要
Self-supervised monocular depth estimation is a critical task in computer vision. Existing methods can be typically categorized into discriminative-based and generative-based methods according to different data modeling approaches. Discriminative-based methods are distinguished by high accuracy, while generative-based methods are notable for superior robustness. Given that images captured in real-world scenarios are inevitably influenced by various external factors, it is essential to develop a robust and accurate depth estimation algorithm. However, there is limited research on the balance of robustness and accuracy by exploring the interactions between discriminative and generative networks. We propose a generative diffusion-based self-supervised monocular depth estimation algorithm guided by discriminative networks and incorporate a depth interaction constraint. We utilize discriminative networks to optimize image-guided information for the denoising process within the diffusion model. This approach seamlessly combines great robustness with high accuracy. Additionally, to reduce the impact of low texture regions on the reprojection photometric loss, we design a texture-aware discriminatory mask module. This module strengthens the constraint capability of the photometric consistency. We conduct experiments on the KITTI and Make3D datasets. The results demonstrate that our method successfully balances accuracy and robustness.