Semantic-Aware Multi-Exposure Fusion Through Vision-Language Model
摘要
The growing demand for High Dynamic Range (HDR) imaging in digital imaging has driven significant advances in Multi-Exposure Image Fusion (MEF), yet existing methods remain constrained by their reliance on low-level visual features while neglecting semantic context. Recent breakthroughs in multimodal large language models like GPT-4o have demonstrated unprecedented cross-modal understanding capabilities, with Vision-Language Model (VLM) such as CLIP establishing new paradigms for semantic-visual correlation. To bridge this semantic-awareness gap in MEF, we propose a novel framework: Semantic-Aware Multi-Exposure Image Fusion via Vision-Language Model (VLM-MEF), that synergizes linguistic guidance with visual processing. Our solution processes luminance and chrominance components through dual pathways, where semantic descriptors generated by vision-language interaction dynamically guide feature fusion. The framework preserves natural color distributions via channel-adaptive recalibration and enforces semantic consistency through text-guided optimization. Extensive validation demonstrates superior performance in structural preservation across diverse exposure scenarios and enhanced color fidelity in saturated chromatic regions compared to conventional approaches, while achieving more visually coherent and perceptually natural fusion results.