<p>Machine learning-based face re-aging enables to automatically change age-related attributes given the target age, starkly reducing the need for manual labor among experienced artists. Variants of StyleGAN or diffusion models have shown promising results, but these approaches often lead to low fidelity on in-the-wild samples and/or still remain in image space. Therefore, their video-level technics have lagged behind, although video inference is imperative for practical aspects. To this end, we introduce diffusion-based re-aging models, marking the first attempt to incorporate diffusion scheme into the video face re-aging task. Our optimization-based denoising approach can produce faithful re-aging results under various conditions, surpassing the shortcoming of GANs and VAEs. Concretely, to improve global semantic coherence, joint null-text optimization is proposed where a single embedding is learned to cover the entire scene using keyframes. In addition, we leverage delta maps which significantly amplify image fidelity through residual manner. Thanks to the nature of delta, our system can lift the 2D diffusion models to video editing by neglecting age-irrelevant regions and by propagating age-relevant pixels to adjacent frames through estimated optical flow without costly computing power. Our diffusion-based delta strategy ensures high fidelity, achieving unprecedented generalization capabilities on in-the-wild cases such as occlusion and accessories and also satisfying industrial demands such as movie trailers, CGI. Video exhibitions and additional results are available on the project page: <a href="https://gh-bumsookim.github.io/VideoTimeTravel/">https://gh-bumsookim.github.io/VideoTimeTravel/</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VideoTimeTravel: high-fidelity face re-aging diffusion models for production video

  • Bumsoo Kim,
  • Yunyoung Nam,
  • Sanghyun Seo

摘要

Machine learning-based face re-aging enables to automatically change age-related attributes given the target age, starkly reducing the need for manual labor among experienced artists. Variants of StyleGAN or diffusion models have shown promising results, but these approaches often lead to low fidelity on in-the-wild samples and/or still remain in image space. Therefore, their video-level technics have lagged behind, although video inference is imperative for practical aspects. To this end, we introduce diffusion-based re-aging models, marking the first attempt to incorporate diffusion scheme into the video face re-aging task. Our optimization-based denoising approach can produce faithful re-aging results under various conditions, surpassing the shortcoming of GANs and VAEs. Concretely, to improve global semantic coherence, joint null-text optimization is proposed where a single embedding is learned to cover the entire scene using keyframes. In addition, we leverage delta maps which significantly amplify image fidelity through residual manner. Thanks to the nature of delta, our system can lift the 2D diffusion models to video editing by neglecting age-irrelevant regions and by propagating age-relevant pixels to adjacent frames through estimated optical flow without costly computing power. Our diffusion-based delta strategy ensures high fidelity, achieving unprecedented generalization capabilities on in-the-wild cases such as occlusion and accessories and also satisfying industrial demands such as movie trailers, CGI. Video exhibitions and additional results are available on the project page: https://gh-bumsookim.github.io/VideoTimeTravel/