A dense triple-level attention-based network for surgical instrument segmentation
摘要
In robot-assisted surgery, precise and autonomous segmentation of surgical instruments is pivotal yet challenged by complexities such as intricate surgical scenarios, variable perspectives, low-contrast imaging, diverse instrument dimensions and forms, alongside class imbalance issue. While deep learning has significantly progressed in pixel-wise image segmentation, prevailing end-to-end models grapple with limitations including inadequate context understanding, inefficient feature extraction, and constrained receptive fields, leading to blurred boundaries and loss of local details. To surmount these challenges and enhance segmentation accuracy, this paper proposes a Dense Triple-Attention Network, named DTA-Net, specifically tailored for pixel-level segmentation of surgical instruments in robot-assisted interventions. Proposed DTA-Net introduces a dense attention block to bolster contextual representation, tackling the deficiency in capturing extensive surgical contexts. In addition, it further integrates a Triple-Attention (TA) block to refine feature processing, thereby extracting intricate pixel relationships and enhancing detail retrieval. Complementarily, a Multi-Scale Context Fusion (MCF) mechanism is devised to consolidate multi-resolution features, ensuring comprehensive context exploitation. Extensive evaluations on benchmark datasets, Kvasir-Instrument and Endovis2017, demonstrate that proposed DTA-Net outperforms state-of-the-art models in terms of segmentation accuracy and completeness, which underscores the potential in advancing the precision of robotic surgical interventions through refined instrument segmentation.