错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Geniebrush: a text-guided multimodal image restoration and editing framework for complex scenes

  • Hongqing Wang,
  • Jun Kit Chaw,
  • Marizuana Mat Daud,
  • Liantao Shi,
  • Nannan Huang,
  • Tin Tin Ting,
  • Yong Wang

摘要

We present GenieBrush, a text-guided multimodal framework for image restoration and editing in complex scenes. Unlike existing models that rely on generic diffusion backbones or static attention mechanisms, GenieBrush integrates language-conditioned reasoning, fine-grained localization, and dynamic refinement to achieve faithful and controllable visual editing. Specifically, we propose three novel modules: (1) Object-Conditioned Alignment Module (OCAM), which decomposes images into slot-based object representations and performs gated alignment with textual instructions; (2) Spatially-Preserving Instruction Network (SPIN), which generates discrete edit masks through vector-quantized alignment of vision-language embeddings, enabling precise target localization; and (3) Noise-Aware Refinement Engine (NARE), which dynamically adjusts the denoising schedule based on structural feedback to preserve fine textures and semantic integrity. Built upon Qwen2.5-VL and Stable Diffusion 1.5 with LoRA-based tuning, GenieBrush achieves state-of-the-art performance on MagicBrush and EMU Edit benchmarks. It surpasses recent baselines by 2.8% CLIP-I, reduces L1/L2 errors by 11.2%/14.1%, and exhibits superior DINO similarity. Extensive ablation studies validate the necessity of each proposed module. Our framework offers a generalizable and scalable solution for open-ended multimodal image editing tasks.