MMTU-Net: enhancing medical image semantic segmentation with multi-level multi-scale fusion and transformer
摘要
Semantic segmentation in medical imaging remains challenging due to issues such as semantic information loss during downsampling, excessive semantic gaps in Skip-connections, and the neglect of global information by deep networks. To address these challenges, we propose MMTU-Net, a U-shaped network augmented with a Multi-level Fusion Transformer (MFT) in the deep layers to expand the receptive field. A dual-channel attention mechanism (DCA) is introduced to reorganize channel weights, preventing insufficient edge information extraction. Furthermore, a Double-Layer Multi-Scale Feature Fusion (DMFF) module is integrated into the Skip-connections to merge images of different levels, thereby reducing feature loss and semantic gaps. Evaluated on four medical image datasets using seven metrics, MMTU-Net demonstrates superior performance compared to benchmark models. Ablation studies confirm the effectiveness of its key components.