<p>Multi-focus image fusion aims to integrate images captured at different focal distances into a single, high-quality image. We proposed a method based on cross-task semantic interaction and dual-attention mixing transformer. Specifically, the dual-attention mixing transformer combines spatial and channel self-attention to capture both local and global information. To overcome common issues of distortion and artifacts, we adopted a multitask learning strategy that simultaneously predicts the decision map and generates the fused image, ensuring natural transitions at focus/defocus boundaries through boundary detection. The cross-task semantic interaction module integrates the spatial and channel information of two tasks. By using a cross-attention transformer, it enhances the complementary information between them, allowing both tasks to learn beneficial semantic information from each other. We evaluated our method against eleven state-of-the-art methods on four multi-focus image datasets. Results show that our approach achieves superior subjective and objective performance while maintaining lightweight design and high efficiency. The source code is available at <a href="https://github.com/ZYZ-GPU/CSI-DMT">https://github.com/ZYZ-GPU/CSI-DMT</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CSI-DMT: multi-focus image fusion via cross-task semantic interaction and dual-attention mixing transformer

  • Hao Zhai,
  • Yuanzhe Zhang,
  • Zhi Zeng,
  • Minyu Deng,
  • Yiyang Ru

摘要

Multi-focus image fusion aims to integrate images captured at different focal distances into a single, high-quality image. We proposed a method based on cross-task semantic interaction and dual-attention mixing transformer. Specifically, the dual-attention mixing transformer combines spatial and channel self-attention to capture both local and global information. To overcome common issues of distortion and artifacts, we adopted a multitask learning strategy that simultaneously predicts the decision map and generates the fused image, ensuring natural transitions at focus/defocus boundaries through boundary detection. The cross-task semantic interaction module integrates the spatial and channel information of two tasks. By using a cross-attention transformer, it enhances the complementary information between them, allowing both tasks to learn beneficial semantic information from each other. We evaluated our method against eleven state-of-the-art methods on four multi-focus image datasets. Results show that our approach achieves superior subjective and objective performance while maintaining lightweight design and high efficiency. The source code is available at https://github.com/ZYZ-GPU/CSI-DMT.