SWSTformer: A stationary wavelet-spatial attention multi-scale transformer framework for single image deraining
摘要
Rain significantly reduces visibility and impairs the performance of sensor-based perception systems in tasks such as autonomous driving and video surveillance. However, most existing methods emphasize spatial or global features while neglecting frequency-domain cues that are crucial for distinguishing rain streaks from background textures. To address this limitation, we propose the Stationary Wavelet-Spatial Transformer (SWSTformer), a novel Transformer-based architecture that jointly models spatial and frequency representations. Each hierarchical stage incorporates a stationary wavelet-spatial transformer block with two key components: a wavelet-spatial window attention mechanism for cross-scale and directional feature modeling, and a dual convolutional gated feed-forward network for local refinement and global context aggregation. To enhance inter-level consistency, we further introduce a dynamic pixelwise cross-adaptive gating network for adaptive feature propagation and a multi-scale adaptive patch embedding module to improve sensitivity to diverse rain patterns. Extensive experiments on both synthetic and real-world datasets demonstrate that SWSTformer achieves state-of-the-art performance, effectively restoring fine details and preserving structural integrity, while enhancing the robustness of sensor-based perception in real-world scenarios.