Can Language Really Understand Depth?
摘要
Vision-Language Pre-training (e.g., CLIP) bridges image and language. Despite its great success in a wide range of zero-shot downstream tasks, it still underperforms in some abstract or systematic tasks, such as classifying the distance to the nearest car. Recently, however, some researchers found that vision-language pre-trained models are able to estimate monocular depth, and their performance even approaches that of some earlier fully-supervised methods. Given these conflicting findings, in this paper, we focus on the question - Can vision-language pre-trained models really understand depth? If so, how well does it perform? To answer these two questions, we propose MonoCLIP, which attempts to fully exploit the potential of vision-language pre-trained models by introducing three basic depth estimators and global context-guided depth fusion. Results on two mainstream monocular depth estimation datasets demonstrate the ability of vision-language pre-trained model in understanding depth. Moreover, adequate ablation studies further shed light on why and how it works.