CMIT-Net: a cross-modal information transfer network for multi-modal brain tumor segmentation
摘要
Automatic segmentation of brain tumor regions from multimodal MRI brain tumor images is essential for clinical diagnosis and treatment. Due to the intricate relationships among various MRI modal images, it has been challenging to effectively utilize the complementary information among different modalities. However, instead of allowing other modal data to take the role of complementary information while extracting information from one modality, most existing works focus on fusion at the bottleneck layer of the model or only fusing multi-modal images into a single-modal image. In this paper, we propose a cross-modal information transfer network for multi-modal brain tumor segmentation (CMIT-Net). Specifically, we present an encoding path for each modal image to extract their features separately and a cross-modal transformer module among different encoding paths, which can utilize the complementary information from other modalities while extracting the information of each modality. At the bottleneck level of the model, we further present a Transformer with three planes module to capture the long-range relationship in each modal image. To better help the decoder receive salient patterns during the upsampling from encoder features, we propose a joint guidance fusion module to efficiently integrate the encoder features and the decoder features. We perform extensive experiments on BraTS2020 challenge dataset and BraTS2021 challenge dataset. Experimental results show that CMIT-Net achieves state-of-the-art segmentation performance compared with existing excellent models.