DUC-Net: A deep unified cross-modal architecture with cascaded feature learning for medical image segmentation
摘要
Accurate segmentation of placental vasculature in fetoscopic imaging is critical for improving intraoperative precision and reducing complications during fetal surgery. However, the low fidelity and complex nature of fetoscopic images pose significant challenges for traditional segmentation methods. This paper presents a novel DUC-Net that effectively balances local detail preservation with global context modeling for computer-assisted fetoscopic surgery. Our primary innovation is the hybrid-stream cross encoder, which integrates a bidirectional dual-encoder mechanism with a specialized dual-pathway attention. This attention enables efficient cross-modal feature integration through a parallel mechanism, enhancing computational efficiency without compromising feature fidelity. We further enhance the architecture with a feature unification module that merges complementary features using adaptive spatial attention and channel recalibration. Furthermore, the decoder network incorporates adaptive learning block to optimize feature reconstruction during upsampling. Extensive experiments on clinical datasets demonstrate that our framework achieves a Dice score of 95.09%, outperforming state-of-the-art methods in vessel boundary delineation and structural connectivity preservation. Comprehensive evaluation across four diverse segmentation tasks validates the superior performance of the proposed method, thereby establishing a new benchmark for computer-assisted medical image segmentation.