Fourier-Enhanced Swin Transformer: An Image Denoising Approach for Intelligent Humanoid Robot Vision Systems
摘要
Recent progress in intelligent humanoid robotics demands advanced visual recognition pipelines. One major obstacle is image noise, which can severely degrade vision-based performance. To address this, we introduce the Fourier-Enhanced Swin Transformer (FESwinT), a dual-phase framework encompassing feature extraction followed by image reconstruction. During extraction, a local attention strategy with a sliding-window approach enables global context modeling and captures rich feature representations, which are then refined via the Fourier transform. These refined representations feed into the reconstruction part to generate the final denoised output, effectively eliminating noise. Thorough experiments on standard denoising benchmarks reveal that FESwinT surpasses current state-of-the-art methods at noise levels σ = 15, 25, and 50 for both grayscale and color images.