Depth estimation in real-world videos has been extensively researched, however, surgical videos pose unique challenges, such as specular reflections, multifaceted occlusions from tissues, fluid and surgical instruments. Accurately estimating depth from a 2D perspective amidst these conditions demands expertise, experience and cognitive effort. Implementing real-time depth estimation especially in endoscopic surgeries, could assist surgeons, leading to a decrease in post-operative complications. In this paper, we present a novel methodology for self-supervised monocular depth estimation using stereo pair images. The contributions of this study are, 1) a modified Siamese(Twin) network encoder-decoder architecture with a Gated Fusion DiffNet using Vision Transformers(DVTwin), 2) utility of a combination loss, fusing Structural Similarity Index Measure and L1 Loss, 3) utility of a random-masking technique. The ViT encoder leverages self-attention and captures global spatial information. The Gated Fusion DiffNet in the encoder calculates disparity at each stage of the network. The combination loss captures structural information from the stereo images while preserving sharp discontinuities at the edges. Training the model with random masking teaches it to learn depth from missing information enabling it to function efficiently with occlusions. We train and evaluate our model on 12 videos of the publicly available Hamlyn dataset. The videos comprise of challenging intra-corporeal scenes from endoscopic surgeries. On the holdout set, we report an Absolute Relative Error(AbsRel) of 0.084 and RMSE of 8.352, a 3% improvement from SOTA. To test for generalizability, we evaluate our model on the test set of SCARED dataset. We achieve an AbsRel of 0.067 and RMSE of 5.953, on par with SOTA, illustrating our model’s generalizability. We fine-tune the model on the train set of SCARED dataset and evaluate on the test set to achieve an improvement of 16% in AbsRel and 18% in RMSE.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-supervised Siamese Network Using Vision Transformer for Depth Estimation in Endoscopic Surgeries

  • Snigdha Agarwal,
  • Neelam Sinha

摘要

Depth estimation in real-world videos has been extensively researched, however, surgical videos pose unique challenges, such as specular reflections, multifaceted occlusions from tissues, fluid and surgical instruments. Accurately estimating depth from a 2D perspective amidst these conditions demands expertise, experience and cognitive effort. Implementing real-time depth estimation especially in endoscopic surgeries, could assist surgeons, leading to a decrease in post-operative complications. In this paper, we present a novel methodology for self-supervised monocular depth estimation using stereo pair images. The contributions of this study are, 1) a modified Siamese(Twin) network encoder-decoder architecture with a Gated Fusion DiffNet using Vision Transformers(DVTwin), 2) utility of a combination loss, fusing Structural Similarity Index Measure and L1 Loss, 3) utility of a random-masking technique. The ViT encoder leverages self-attention and captures global spatial information. The Gated Fusion DiffNet in the encoder calculates disparity at each stage of the network. The combination loss captures structural information from the stereo images while preserving sharp discontinuities at the edges. Training the model with random masking teaches it to learn depth from missing information enabling it to function efficiently with occlusions. We train and evaluate our model on 12 videos of the publicly available Hamlyn dataset. The videos comprise of challenging intra-corporeal scenes from endoscopic surgeries. On the holdout set, we report an Absolute Relative Error(AbsRel) of 0.084 and RMSE of 8.352, a 3% improvement from SOTA. To test for generalizability, we evaluate our model on the test set of SCARED dataset. We achieve an AbsRel of 0.067 and RMSE of 5.953, on par with SOTA, illustrating our model’s generalizability. We fine-tune the model on the train set of SCARED dataset and evaluate on the test set to achieve an improvement of 16% in AbsRel and 18% in RMSE.