Transformer-cnn hybrid network for underwater image enhancement
摘要
Due to light attenuation and wavelength-dependent light scattering in underwater environments, underwater images often suffer from color distortion, reduced contrast, and loss of detail, which significantly hinder the performance of high-level computer vision tasks. To address these challenges, we propose a hybrid Transformer-CNN network for underwater image enhancement. First, a Mixed Convolution and Transformer Block (MCTB) is employed in both the encoder and decoder, integrating hierarchical convolutions with self-attention mechanisms to enhance the model’s ability to capture fine-grained details and global contextual information. Secondly, to mitigate feature interference caused by the symmetrical structure of the encoder-decoder architecture, an Enhanced Feature Fusion Module (EFFM) is introduced in the skip connections to refine color and structural information. Finally, the model is trained on the publicly available UIEB-paired dataset and evaluated on multiple underwater image test sets. Quantitative and qualitative results demonstrate that the proposed method achieves superior performance compared to other methods.