Deep Learning for Multimodal Medical Image Fusion: A Concise Review
摘要
Multimodal medical image fusion combines images from different scanners so one image carries more useful and complementary information. Modern deep learning methods—such as convolutional neural networks (CNNs), generative adversarial networks (GANs), and transformers—learn how to fuse images directly from data. In this concise review, we explain the main ideas, summarize representative model families, describe common attention and feature enhancement modules, outline training strategies, and list typical evaluation metrics and clinical uses. We also point out practical issues: limited labeled data, different scanners and sites, the need for interpretability, and compute limits in clinics. Finally, we highlight priorities for future work, including self- supervised learning, explainable and trustworthy modules, and efficient architectures that are easier to deploy in practice.