Medical image fusion (MIF) aims to integrate complementary information from multimodal images, generating a single fused image for enhanced applications in clinical diagnosis, preoperative planning. However, existing MIF models primarily focus on modal complementarity, often neglecting the problem of modal misalignment at different scales. This oversight can result in an over-reliance on a single modality in the fused image, thereby compromising the balanced fusion of multimodal information. Therefore, we proposed a Learnable Modal Alignment Network (LMAFusion) for the MIF task. Specifically, to alleviate the modal misalignment problem, we develop a Learnable Modal Alignment (LMA) module, which computes the maximum value, mean, and standard deviation of each modality’s features. These statistics are fed into a Multi-Layer Perception (MLP) to automatically estimate the optimal parameters, guiding the heterogeneous modalities to adaptively adjust their features for alignment. In addition, to prevent a single modality from dominating the fusion result, we adopt a cyclic iterative strategy that dynamically adjusts the input modal features of the LMA module, promoting balanced expression of heterogeneous modal features. Furthermore, a loss function consisting of the sum of the correlation of differences loss and structural loss drives the LMAFusion network to preserve rich complementary modal information. Extensive experiments on SPECT-MRI and PET-MRI datasets show that our LMAFusion model outperforms 15 State-of-the-art (SOTA) image fusion algorithms in both subjective visual assessment and objective metric evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rethinking the Necessity of Learnable Modal Alignment for Medical Image Fusion

  • Min Li,
  • Feng Li,
  • Enguang Zuo,
  • Xiaoyi Lv,
  • Chen Chen,
  • Cheng Chen

摘要

Medical image fusion (MIF) aims to integrate complementary information from multimodal images, generating a single fused image for enhanced applications in clinical diagnosis, preoperative planning. However, existing MIF models primarily focus on modal complementarity, often neglecting the problem of modal misalignment at different scales. This oversight can result in an over-reliance on a single modality in the fused image, thereby compromising the balanced fusion of multimodal information. Therefore, we proposed a Learnable Modal Alignment Network (LMAFusion) for the MIF task. Specifically, to alleviate the modal misalignment problem, we develop a Learnable Modal Alignment (LMA) module, which computes the maximum value, mean, and standard deviation of each modality’s features. These statistics are fed into a Multi-Layer Perception (MLP) to automatically estimate the optimal parameters, guiding the heterogeneous modalities to adaptively adjust their features for alignment. In addition, to prevent a single modality from dominating the fusion result, we adopt a cyclic iterative strategy that dynamically adjusts the input modal features of the LMA module, promoting balanced expression of heterogeneous modal features. Furthermore, a loss function consisting of the sum of the correlation of differences loss and structural loss drives the LMAFusion network to preserve rich complementary modal information. Extensive experiments on SPECT-MRI and PET-MRI datasets show that our LMAFusion model outperforms 15 State-of-the-art (SOTA) image fusion algorithms in both subjective visual assessment and objective metric evaluation.