错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MAGIC: Multi-prompt Any Length Video Generation Model with Controllable Inter-frame Correlation and Low Barrier

  • Jialiang Xu,
  • Weiran Chen,
  • Lingbing Xu,
  • Weitao Song,
  • Yi Ji,
  • Ying Li,
  • Chunping Liu

摘要

For the text-to-video (T2V) task, high inter-frame correlation and low barrier for potential users are both desired and tackled by most existing models. Therefore, in this paper, we propose a novel T2V model based on the text-to-image (T2I) diffusion model, called MAGIC. It solves the problem of inter-frame correlation, and the correlation can be controlled. Moreover, it can generate videos at a low cost. In addition, MAGIC can also support multi-prompt (text), and multi-image input for content control, which generates videos of any length with a diversity of content. Specifically, we integrate noise, condition, and adapter information. This strategy allows MAGIC to generate high-quality video with better continuity and controllability of neighboring frames. Furthermore, our approach does not require training. It can run on just one GPU and only adds about 0.0015% GLOPs computational overhead based on the T2I model, which significantly lowers the barrier of computational resources. Experiments show that MAGIC can generate high-quality videos with better inter-frame correlation and higher-quality video content. Combining the advantages of being able to freely input multi-prompt and multi-image, our approach can not only make generating videos more controllable but also meet individual needs. Our MAGIC demonstrates significant potential in the field of video generation.