<p>Multimodal data fusion is pivotal in artificial intelligence, with infrared and visible image fusion crucial for enhancing target detection. However, low-light visible images often lack vital details, impacting fusion quality and target detection accuracy, particularly in autonomous driving applications. To address these challenges, an innovative model is proposed that introduces a normalized depth-wise crested porcupine cross-attention model to extract complementary details from visible and infrared images. The incorporation of double cross-attention effectively enhances complementary information while minimizing redundant features. Subsequently, the Decomposed Non-subsampled Shearlet Structural Patch Transform (DNSSPT) model is employed to integrate features and generate the fused image. This process involves decomposing the extracted infrared and visible images into components such as signal intensity, mean intensity, and signal structure through structural patch decomposition. Next, a membership curve is applied to precisely determine the weight of the average intensity module, minimizing artifacts while preserving the importance of infrared targets. Additionally, sharpening operations are used to improve detail layer of both the visible and infrared images, leading to a fused image with higher contrast. Through subjective and objective evaluations, the proposed model outperforms existing fusion techniques, achieving minimum processing times of 0.0521, 0.0698, 0.0681, and 0.0832&#xa0;s across the TNO, MSRS, VIFB, and VOT-RGBT datasets, respectively. Additionally, it attains high detection accuracies of 94% for VOT-RGBT, 92.71% for TNO, 96% for MSRS, and 90.89% for VIFB. Furthermore, object detection experiments using roadscene dataset confirm the DNSSPT model’s effectiveness in advancing computer vision tasks, achieving an accuracy of 94.1%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Fusion of Infrared and Visible Images Using an Optimized Deep Learning Model with Decomposed Non-subsampled Shearlet Structural Patch Transform

  • P. Murugeswari,
  • Sonia Jenifer Rayen,
  • M. Ramkumar,
  • U. Sakthi

摘要

Multimodal data fusion is pivotal in artificial intelligence, with infrared and visible image fusion crucial for enhancing target detection. However, low-light visible images often lack vital details, impacting fusion quality and target detection accuracy, particularly in autonomous driving applications. To address these challenges, an innovative model is proposed that introduces a normalized depth-wise crested porcupine cross-attention model to extract complementary details from visible and infrared images. The incorporation of double cross-attention effectively enhances complementary information while minimizing redundant features. Subsequently, the Decomposed Non-subsampled Shearlet Structural Patch Transform (DNSSPT) model is employed to integrate features and generate the fused image. This process involves decomposing the extracted infrared and visible images into components such as signal intensity, mean intensity, and signal structure through structural patch decomposition. Next, a membership curve is applied to precisely determine the weight of the average intensity module, minimizing artifacts while preserving the importance of infrared targets. Additionally, sharpening operations are used to improve detail layer of both the visible and infrared images, leading to a fused image with higher contrast. Through subjective and objective evaluations, the proposed model outperforms existing fusion techniques, achieving minimum processing times of 0.0521, 0.0698, 0.0681, and 0.0832 s across the TNO, MSRS, VIFB, and VOT-RGBT datasets, respectively. Additionally, it attains high detection accuracies of 94% for VOT-RGBT, 92.71% for TNO, 96% for MSRS, and 90.89% for VIFB. Furthermore, object detection experiments using roadscene dataset confirm the DNSSPT model’s effectiveness in advancing computer vision tasks, achieving an accuracy of 94.1%.