Recent progress in intelligent humanoid robotics demands advanced visual recognition pipelines. One major obstacle is image noise, which can severely degrade vision-based performance. To address this, we introduce the Fourier-Enhanced Swin Transformer (FESwinT), a dual-phase framework encompassing feature extraction followed by image reconstruction. During extraction, a local attention strategy with a sliding-window approach enables global context modeling and captures rich feature representations, which are then refined via the Fourier transform. These refined representations feed into the reconstruction part to generate the final denoised output, effectively eliminating noise. Thorough experiments on standard denoising benchmarks reveal that FESwinT surpasses current state-of-the-art methods at noise levels σ = 15, 25, and 50 for both grayscale and color images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fourier-Enhanced Swin Transformer: An Image Denoising Approach for Intelligent Humanoid Robot Vision Systems

  • Jinwen Niu,
  • Yan Ma,
  • Liang He,
  • Shengjie Guo,
  • Hao Sun

摘要

Recent progress in intelligent humanoid robotics demands advanced visual recognition pipelines. One major obstacle is image noise, which can severely degrade vision-based performance. To address this, we introduce the Fourier-Enhanced Swin Transformer (FESwinT), a dual-phase framework encompassing feature extraction followed by image reconstruction. During extraction, a local attention strategy with a sliding-window approach enables global context modeling and captures rich feature representations, which are then refined via the Fourier transform. These refined representations feed into the reconstruction part to generate the final denoised output, effectively eliminating noise. Thorough experiments on standard denoising benchmarks reveal that FESwinT surpasses current state-of-the-art methods at noise levels σ = 15, 25, and 50 for both grayscale and color images.