<p>Image fusion in the healthcare domain has potential to greatly transform patient treatment, medication planning, and health diagnostics by combining important data from multiple imaging techniques. Current image fusion techniques demonstrate promising outcomes; however, they often overlook the significance of long-range dependencies and inter-scale information within images, primarily relying on manual extraction of image characteristics. This may result in compromised fusion performance. To address this shortcoming, a more effective and robust automated system is proposed to extract the overall image characteristics and deal with long range dependencies. The proposed model, M3IFT2.0 employs a transformer-based framework to autonomously derive characteristics of medical images. The method also incorporates a distinctive attention-based two-dimensional fusion strategy that improves the efficiency and resilience of medical image fusion through the utilization of L2 norm in both horizontal and vertical dimensions. The proposed model is trained on the publicly accessible Lung-PET-CT-Dx and The Whole Brain Atlas datasets. It is assessed and contrasted with a transformer model trained on the publicly available MS-COCO dataset and a specific medical dataset encompassing several visual modalities to execute image fusion. The evaluation of the fused output is conducted using 13 performance metrics. The results demonstrated its superior visual and quantitative performance across various parameters, including structural similarity, mutual information, entropy, multi-scale similarity, mean square error, root mean square error, and peak signal-to-noise ratio, when compared to state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

M3IFT2.0: multi-modal medical image fusion framework with vision transformer features and attention integration

  • Disha Mohini Pathak,
  • Tribhuwan Kumar Tewari

摘要

Image fusion in the healthcare domain has potential to greatly transform patient treatment, medication planning, and health diagnostics by combining important data from multiple imaging techniques. Current image fusion techniques demonstrate promising outcomes; however, they often overlook the significance of long-range dependencies and inter-scale information within images, primarily relying on manual extraction of image characteristics. This may result in compromised fusion performance. To address this shortcoming, a more effective and robust automated system is proposed to extract the overall image characteristics and deal with long range dependencies. The proposed model, M3IFT2.0 employs a transformer-based framework to autonomously derive characteristics of medical images. The method also incorporates a distinctive attention-based two-dimensional fusion strategy that improves the efficiency and resilience of medical image fusion through the utilization of L2 norm in both horizontal and vertical dimensions. The proposed model is trained on the publicly accessible Lung-PET-CT-Dx and The Whole Brain Atlas datasets. It is assessed and contrasted with a transformer model trained on the publicly available MS-COCO dataset and a specific medical dataset encompassing several visual modalities to execute image fusion. The evaluation of the fused output is conducted using 13 performance metrics. The results demonstrated its superior visual and quantitative performance across various parameters, including structural similarity, mutual information, entropy, multi-scale similarity, mean square error, root mean square error, and peak signal-to-noise ratio, when compared to state-of-the-art methods.