错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MDA-ViT: Multimodal image fusion using dual attention vision transformer

  • Shrida Kalamkar,
  • Geetha Mary Amalanathan

摘要

Multimodal image fusion is critical in computer vision, enhancing image quality by combining information from diverse modalities. Traditional fusion methods often rely on manual feature extraction, which may not suit heterogeneous data well. Deep learning has shown promise, but challenges persist in effectively fusing multimodal data. To address this, we propose MDA-ViT, a dual attention vision transformer. MDA-ViT leverages self- and multimodal attention mechanisms to selectively fuse features from different modalities, enhancing semantic understanding and adaptability. The model integrates intra- and inter-modal dependencies, preserving spatial structures and semantic relationships. Specifically, MDA-ViT incorporates two distinct attention mechanisms: intra-modal attention captures spatial relationships within individual modalities, while inter-modal attention learns cross-modal interactions. Additionally, a modified self-attention mechanism dynamically adjusts attention weights based on input data characteristics, enhancing model flexibility. Experimental evaluation on benchmark datasets, including the TNO and Harvard Atlas medical datasets, demonstrates MDA-ViT’s superior fusion performance compared to existing methods, showcasing its potential for diverse applications in medical imaging and beyond.