Self-supervised Cascade Training for Monocular Endoscopic Dense Depth Recovery
摘要
Dense depth prediction for 3-D reconstruction of monocular endoscopic images is an essential way to expand the surgical field and augment the perception of surgeons in robotic endoscopic surgery. However, it is generally challenging to precisely estimate the monocular dense depth and reconstruct such a field due to complex surgical fields with a limited field of view, illumination variations, and weak texture information. This work proposes a new framework of self-supervised learning with a two-stage cascade training strategy for dense depth recovery of monocular endoscopic images. While the first stage is to train an initial deep-learning model through sparse depth consistency supervision, the second stage introduces photometric consistency supervision to further train and refine the initial model for improving its capability. Our framework was evaluated on patient data of monocular endoscopic images acquired from colonoscopic procedures, with the experimental results demonstrating that our self-supervised learning model with cascade training provides a promising strategy outperforming other models. On the one hand, both visual quality and quantitative assessment of our method are better than current monocular dense depth estimation approaches. On the other hand, our method relies less on sparse depth data for supervision than other self-supervised methods.