Backlit photography encounters challenges such as poor contrast, reduced visibility, and high noise levels, which adversely impact both human perception and computer vision systems. Current deep learning-based enhancement methods typically rely on paired images for training, yet the difficulty in acquiring such pairs gives rise to authenticity issues and artifacts in the enhanced images. Moreover, defining a standard for the ideal enhanced image is problematic due to the inherently subjective nature of lighting conditions. Although previous methods have successfully addressed unsupervised backlit image enhancement, the training process remains complex, with significant potential for improvement in both methodology and performance. In this regard, we propose a novel unsupervised approach that leverages the visual-semantic priors of the CLIP model to guide the enhancement process. Specifically, we begin by fine-tuning the CLIP image encoder using unpaired image data, refining its visual priors for backlit image enhancement. Subsequently, we achieve effective unsupervised backlit image enhancement through the use of the specialized CLIP image encoder tailored for this task, coupled with the CLIP text encoder, which provides integrated visual and semantic supervision. Our approach outperforms existing methods on the BAID dataset across multiple quality assessment metrics, establishing a new state-of-the-art for unsupervised backlit image enhancement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative Adversarial CLIPs for Unsupervised Backlit Image Enhancement

  • Dong Tang,
  • Suoyang Sun,
  • Yujun Huang,
  • Bin Chen

摘要

Backlit photography encounters challenges such as poor contrast, reduced visibility, and high noise levels, which adversely impact both human perception and computer vision systems. Current deep learning-based enhancement methods typically rely on paired images for training, yet the difficulty in acquiring such pairs gives rise to authenticity issues and artifacts in the enhanced images. Moreover, defining a standard for the ideal enhanced image is problematic due to the inherently subjective nature of lighting conditions. Although previous methods have successfully addressed unsupervised backlit image enhancement, the training process remains complex, with significant potential for improvement in both methodology and performance. In this regard, we propose a novel unsupervised approach that leverages the visual-semantic priors of the CLIP model to guide the enhancement process. Specifically, we begin by fine-tuning the CLIP image encoder using unpaired image data, refining its visual priors for backlit image enhancement. Subsequently, we achieve effective unsupervised backlit image enhancement through the use of the specialized CLIP image encoder tailored for this task, coupled with the CLIP text encoder, which provides integrated visual and semantic supervision. Our approach outperforms existing methods on the BAID dataset across multiple quality assessment metrics, establishing a new state-of-the-art for unsupervised backlit image enhancement.