Self-supervised monocular depth estimation via multiple bilateral consistency
摘要
Recent researches on deep-learning methods have shown enormous promise for depth estimates. Although self-supervised estimators alleviate laborious annotation, numerous pixels fail to match the corresponding ones in adjacent frames. Single re-projection loss functions are oversimplified for stereoscopic adjacent frame constraints. In this work, we constrain the primitive inter-frame-supervised depth estimation via multiple bilateral consistency, which builds a bi-directional mapping between adjacent frames with inherent properties, allowing to develop pose-consistent and depth-consistent models. For the mismatching pixels in adjacent frames, a cycle-consistency framework reconstructs depth maps as scene images with an additional symmetrical re-rendering network. Pose-consistent constraint aims to ensure the reversibility of ego-motion transformations between adjacent frames, and depth-consistent constraint strives to ensure the continuity of adjacent frames’ depths. With the joint optimization of three independent modules, depth network, pose network and re-rendering network, the proposed framework yields state-of-the-art metrics on the KITTI depth dataset pre-trained on CityScapes, with the AbsRel, SqRel, RMSE and RMS(log) decreased by 6.6%, 12.4%, 0.4%, and 2.2%, respectively.