Enhancing Semantic-Guided Self-supervised Monocular Depth Estimation by Exploring Task-Related Representations
摘要
The semantic-guided depth estimation approach can better understand scene information than traditional methods. These methods utilize a shared encoder followed by a two-branch decoder architecture, simplifying the network but failing to capture task-specific details in high-level features. Additionally, some works introduce skip connections or information interaction mechanisms in the decoder to enrich features. However, static feature fusion methods do not adapt well to varying scene changes. To tackle these issues, first, we propose a multi-scale feature refinement mechanism to refine the features of the encoder. Second, we design a dynamic perception fusion decoder that can adjust adaptively during the feature fusion process. Results from experiments show that our approach produces depth maps of excellent quality and continues to perform better in challenging settings.