<p>Foggy conditions pose significant challenges for the perception systems of autonomous vehicles, as they reduce visibility and compromise safety. While traditional dehazing methods often struggle to generalize across varying fog densities, transformer-based architectures have shown promise yet remain underexplored for this task. In this study, we propose a novel noise-learning paradigm in which Vision Transformers are trained to estimate fog maps rather than directly predict clean images. We apply this approach across three transformer-based architectures: Masked Autoencoders (MAE), Convolutional Vision Transformers (CvT), and Swin Transformers, conducting a comprehensive comparative evaluation. Our framework involves two stages: first, training each architecture using standard direct dehazing; and second, retraining them using the proposed noise-learning strategy. Extensive experiments on both synthetic (RESIDE) and real-world (Foggy Cityscapes) datasets show that noise learning consistently enhances performance across all models, with Swin Transformer achieving the highest gains (34.62 dB PSNR, 0.981 SSIM). These results demonstrate that our approach not only improves dehazing quality but also enables stronger generalization to diverse fog conditions. Moreover, the Swin Transformer stands out as the most promising candidate for real-time autonomous driving due to its effective balance between accuracy and computational efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Noise-learning transformers for image dehazing

  • Mostafa Elgendy,
  • Shimaa Ragab,
  • Ahmed Shalaby

摘要

Foggy conditions pose significant challenges for the perception systems of autonomous vehicles, as they reduce visibility and compromise safety. While traditional dehazing methods often struggle to generalize across varying fog densities, transformer-based architectures have shown promise yet remain underexplored for this task. In this study, we propose a novel noise-learning paradigm in which Vision Transformers are trained to estimate fog maps rather than directly predict clean images. We apply this approach across three transformer-based architectures: Masked Autoencoders (MAE), Convolutional Vision Transformers (CvT), and Swin Transformers, conducting a comprehensive comparative evaluation. Our framework involves two stages: first, training each architecture using standard direct dehazing; and second, retraining them using the proposed noise-learning strategy. Extensive experiments on both synthetic (RESIDE) and real-world (Foggy Cityscapes) datasets show that noise learning consistently enhances performance across all models, with Swin Transformer achieving the highest gains (34.62 dB PSNR, 0.981 SSIM). These results demonstrate that our approach not only improves dehazing quality but also enables stronger generalization to diverse fog conditions. Moreover, the Swin Transformer stands out as the most promising candidate for real-time autonomous driving due to its effective balance between accuracy and computational efficiency.