Monocular Depth Estimation (MDE) is an inherently ill-posed problem due to the lack of binocular depth cues, despite this there have been significant research done in this field in recent years. In an attempt to bridge understanding between human and machine perception, this paper investigates learned concepts from the general-purpose model Depth Anything, focusing on features that are known to be present in the human visual system. We perform interventions on different image features within the KITTI and NYUv2 dataset, evaluating performance on these intervened inputs. This led to interesting insights on how and how much each of these features influence depth perception. These insights contribute to bridging understanding of how humans and machines perform MDE respectively, and we also hope it provides a new way for future work to devise more robust methods of training neural networks for MDE.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature Contribution in Monocular Depth Estimation

  • Hui Yu Lau,
  • Srinandan Dasmahapatra,
  • Hansung Kim

摘要

Monocular Depth Estimation (MDE) is an inherently ill-posed problem due to the lack of binocular depth cues, despite this there have been significant research done in this field in recent years. In an attempt to bridge understanding between human and machine perception, this paper investigates learned concepts from the general-purpose model Depth Anything, focusing on features that are known to be present in the human visual system. We perform interventions on different image features within the KITTI and NYUv2 dataset, evaluating performance on these intervened inputs. This led to interesting insights on how and how much each of these features influence depth perception. These insights contribute to bridging understanding of how humans and machines perform MDE respectively, and we also hope it provides a new way for future work to devise more robust methods of training neural networks for MDE.