Gradient-Guided Transformer Network for Power Equipment Infrared and Visible Image Fusion
摘要
Infrared and visible image fusion is essential for intelligent diagnosis systems in power equipment. The adoption of deep learning technology in this field has been prevalent due to its strong representation learning ability. However, existing methods often require numerous iterations of fine-tuning to identify an optimal fusion network. In this paper, we propose a gradient-guided Transformer network for infrared and visible image fusion, which enhances reliability and interpretability in rendering comprehensive images. Specifically, we combine the Sobel operator with the Swin-Transformer architecture to build a gradient residual Transformer module (GRTM). Firstly, GRTM utilizes the Sobel operator to extract structural priors and generate gradient features. Secondly, these gradient maps are supplemented into the attention mechanism to facilitate fine-grained extraction of textural details and salient objects. Extensive experiments conducted on two public datasets with various scenarios demonstrate that the proposed method outperforms state-of-the-art algorithms in both qualitative and quantitative evaluations, highlighting its superior performance and robustness.