Geniebrush: a text-guided multimodal image restoration and editing framework for complex scenes
摘要
We present GenieBrush, a text-guided multimodal framework for image restoration and editing in complex scenes. Unlike existing models that rely on generic diffusion backbones or static attention mechanisms, GenieBrush integrates language-conditioned reasoning, fine-grained localization, and dynamic refinement to achieve faithful and controllable visual editing. Specifically, we propose three novel modules: (1) Object-Conditioned Alignment Module (OCAM), which decomposes images into slot-based object representations and performs gated alignment with textual instructions; (2) Spatially-Preserving Instruction Network (SPIN), which generates discrete edit masks through vector-quantized alignment of vision-language embeddings, enabling precise target localization; and (3) Noise-Aware Refinement Engine (NARE), which dynamically adjusts the denoising schedule based on structural feedback to preserve fine textures and semantic integrity. Built upon Qwen2.5-VL and Stable Diffusion 1.5 with LoRA-based tuning, GenieBrush achieves state-of-the-art performance on MagicBrush and EMU Edit benchmarks. It surpasses recent baselines by 2.8% CLIP-I, reduces L1/L2 errors by 11.2%/14.1%, and exhibits superior DINO similarity. Extensive ablation studies validate the necessity of each proposed module. Our framework offers a generalizable and scalable solution for open-ended multimodal image editing tasks.