COLORSHOP: Color Manipulation of Objects in Videos Using Diffusion Models
摘要
Diffusion models have unlocked unprecedented capabilities in image generation, while their video counterparts still lag behind due to the excessive training cost of temporal modeling. Besides the training burden, generated videos also suffer from issues of inconsistent appearance and structural flickering. To tackle these challenges, we have designed an optimization-free and zero fine-tuning framework called COLORSHOP to implementing editing of the appearance color of objects in a video based on the continuity of VAE in the latent space. When processing each frame, we introduce Foreground diffusion to accelerate the operation speed. During the generation process, we further propose Cross-Frame spatial feature fusion to enhance foreground continuity across frames. Experimental results have shown that, by combining the currently popular diffusion-based image editing algorithm, COLORSHOP has been proven successful in video editing tasks, demonstrating excellent performance in terms of consistency and quality.