Due to the inconsistent exposure across different regions in low light images, existing methods may struggle to balance under-/over- exposed regions. In this paper, we propose a novel Vision Language Model (VLM) guided framework that leverages the inherent semantic understanding and zero-shot reasoning capabilities of Qwen2.5-VL to establish adaptive enhancement guidance. Specifically, we generated dual attention maps utilizing visual prompts to enable an explicit modeling and guidance of low light image enhancement processes. These dual attention maps include contrast attention maps that quantify local exposure and texture attention maps that evaluate structural degradation. Proposed architecture employs a semi-supervised enhancement module (EM) with mamba-based state space modeling, integrating contrast guidance in shallow layers for global adjustment and texture guidance in deeper layers for detail refinement. Moreover, the adaptive guidance module (AGM) dynamically modulates feature processing through self-attention mechanisms conditioned on the VLM-generated guide maps. The training paradigm combines unsupervised pre-training, which is based on color constancy, structural consistency, and exposure constraints, followed by few-shot supervised fine-tuning using diverse illumination samples. Experiments demonstrate that the proposed dual guidance approach effectively coordinates global contrast restoration and local texture enhancement while maintaining natural visual balance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual Attention Guidance with Vision-Language Models for Exposure-Consistent Illumination Enhancement

  • Haodian Wang,
  • Lilin Sui

摘要

Due to the inconsistent exposure across different regions in low light images, existing methods may struggle to balance under-/over- exposed regions. In this paper, we propose a novel Vision Language Model (VLM) guided framework that leverages the inherent semantic understanding and zero-shot reasoning capabilities of Qwen2.5-VL to establish adaptive enhancement guidance. Specifically, we generated dual attention maps utilizing visual prompts to enable an explicit modeling and guidance of low light image enhancement processes. These dual attention maps include contrast attention maps that quantify local exposure and texture attention maps that evaluate structural degradation. Proposed architecture employs a semi-supervised enhancement module (EM) with mamba-based state space modeling, integrating contrast guidance in shallow layers for global adjustment and texture guidance in deeper layers for detail refinement. Moreover, the adaptive guidance module (AGM) dynamically modulates feature processing through self-attention mechanisms conditioned on the VLM-generated guide maps. The training paradigm combines unsupervised pre-training, which is based on color constancy, structural consistency, and exposure constraints, followed by few-shot supervised fine-tuning using diverse illumination samples. Experiments demonstrate that the proposed dual guidance approach effectively coordinates global contrast restoration and local texture enhancement while maintaining natural visual balance.