M3Diff:Semantic Mask-Guided 3D Medical Image Synthesis via Mamba-U-Net Hybrid for Data Augmentation
摘要
Diffusion models have emerged as a promising approach for addressing the scarcity of medical imaging data due to their exceptional generative capabilities. However, current diffusion models face significant limitations in 3D medical image synthesis. First, substantial variations in lesion sizes hinder effective multi-scale feature fusion, while the large volumetric data of 3D images further increases the computational cost of global feature modeling. To address these issues, we propose M3Diff, a novel 3D medical image generation method that enables voxel-level control through semantic segmentation masks. The core innovation of M3Diff lies in the design of a multi-scale spatial-channel feature modulation mechanism, which effectively integrates multi-scale features through dynamic feature interaction. Furthermore, a dual attention adaptive fusion strategy is developed to adaptively merge deep semantic features with shallow high-resolution details. To achieve efficient local feature extraction and global feature integration, we strategically combine Mamba’s efficient long-sequence modeling with U-Net’s robust feature extraction capabilities. Experiments demonstrate that M3Diff achieves state-of-the-art FID scores of 0.184 on BraTS21 and 1.60 on ATLAS v2.0. Moreover, when the synthesized datasets are applied to downstream tasks, segmentation models trained with M3Diff-generated data exhibit Dice score improvements of 2.4 percentage points (pp) on BraTS21 and 3.1 pp on ATLAS v2.0. These results validate M3Diff’s effectiveness in enhancing 3D medical image quality and alleviating data scarcity challenges.