Visible-infrared image fusion has attracted great attention in a range of computer vision applications. Aiming at improving task-specific performance, recent studies have employed a cascading approach, where the fusion network is trained using feedback from the specific downstream task network. However, this training strategy will result in the overfitting of the fusion network, and deploying a different fusion network for each downstream task is inefficient for multi-task scenarios. To address this challenge, we propose VIFA, a visible-infrared image fusion architecture for multi-task applications. This architecture effectively mitigates the catastrophic forgetting problem by partitioning the fusion network into a knowledge-sharing backbone and task-specific components. To facilitate knowledge sharing, we introduce a key channel-constrained distillation strategy, which identifies and retains informative features, while allowing non-critical channels to learn new knowledge. In addition, we propose a reference model-guided distillation to compress the task-specific components while maintaining model performance. Evaluations on multiple representative fusion networks show that VIFA can significantly improve task performance and speed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VIFA: An Efficient Visible and Infrared Image Fusion Architecture for Multi-task Applications via Continual Learning

  • Jiaxing Shi,
  • Ao Ren,
  • Wei Zhuang,
  • Yang Hua,
  • ZhiYong Qin,
  • Zhenyu Wang,
  • Yang Song,
  • Yujuan Tan,
  • Duo Liu

摘要

Visible-infrared image fusion has attracted great attention in a range of computer vision applications. Aiming at improving task-specific performance, recent studies have employed a cascading approach, where the fusion network is trained using feedback from the specific downstream task network. However, this training strategy will result in the overfitting of the fusion network, and deploying a different fusion network for each downstream task is inefficient for multi-task scenarios. To address this challenge, we propose VIFA, a visible-infrared image fusion architecture for multi-task applications. This architecture effectively mitigates the catastrophic forgetting problem by partitioning the fusion network into a knowledge-sharing backbone and task-specific components. To facilitate knowledge sharing, we introduce a key channel-constrained distillation strategy, which identifies and retains informative features, while allowing non-critical channels to learn new knowledge. In addition, we propose a reference model-guided distillation to compress the task-specific components while maintaining model performance. Evaluations on multiple representative fusion networks show that VIFA can significantly improve task performance and speed.