<p>Depth estimation from 2D images is an essential task in computer vision with applications in scene understanding, robotics, and autonomous systems. The performance of supervised depth models depends on network design, loss formulation, data quality, and fine-tuning strategy. In this study, we propose a progressive fine-tuning approach for metric (absolute-scale) depth estimation. Our method uses transfer learning across multiple indoor datasets: real, synthetic, and pseudo-labelled. DenseNet-169 and EfficientNet-B0 backbones are fine-tuned on MIT-G, SUN-RGBD, SceneNet, and NYU2. We apply a three-scale combined loss with weighted MAE + Edge + SSIM terms at full, 1/2, and 1/4 resolution, and add a perceptual VGG component, while we keep the global coefficients of the loss at 1 for simplicity and reproducibility. We find that EfficientNet performs better on the smaller datasets, while DenseNet benefits most from the million-image SceneNet stage and reaches REL 0.105 and RMSE 0.359 on NYU2, comparable to recent transformer baselines yet using 6<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4496_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> fewer parameters. The pseudo-labelled MIT-G data is used as a warm-start and shows the potential of reducing annotation cost. All headline metrics results are based on sensor ground-truth data, avoiding circular evaluation. Qualitative analysis and zero-shot tests on the unseen iBims-1 benchmark confirm that the models generalise and produce coherent, detailed depth maps across diverse indoor scenes. The proposed pipeline thus offers a balanced trade-off between accuracy and computational cost for practical indoor depth estimation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-Source Depth Estimation: Utilizing Real, Synthetic, and Monocular Depth Data with Custom Loss Functions

  • Muhammad Adeel Hafeez,
  • Ganesh Sistu,
  • Michael G. Madden,
  • Ihsan Ullah

摘要

Depth estimation from 2D images is an essential task in computer vision with applications in scene understanding, robotics, and autonomous systems. The performance of supervised depth models depends on network design, loss formulation, data quality, and fine-tuning strategy. In this study, we propose a progressive fine-tuning approach for metric (absolute-scale) depth estimation. Our method uses transfer learning across multiple indoor datasets: real, synthetic, and pseudo-labelled. DenseNet-169 and EfficientNet-B0 backbones are fine-tuned on MIT-G, SUN-RGBD, SceneNet, and NYU2. We apply a three-scale combined loss with weighted MAE + Edge + SSIM terms at full, 1/2, and 1/4 resolution, and add a perceptual VGG component, while we keep the global coefficients of the loss at 1 for simplicity and reproducibility. We find that EfficientNet performs better on the smaller datasets, while DenseNet benefits most from the million-image SceneNet stage and reaches REL 0.105 and RMSE 0.359 on NYU2, comparable to recent transformer baselines yet using 6 \(\times \) × fewer parameters. The pseudo-labelled MIT-G data is used as a warm-start and shows the potential of reducing annotation cost. All headline metrics results are based on sensor ground-truth data, avoiding circular evaluation. Qualitative analysis and zero-shot tests on the unseen iBims-1 benchmark confirm that the models generalise and produce coherent, detailed depth maps across diverse indoor scenes. The proposed pipeline thus offers a balanced trade-off between accuracy and computational cost for practical indoor depth estimation.