The generation of video from textual descriptions is an emerging interdisciplinary research area, merging advancements in natural language processing (NLP), computer vision, and deep learning. This paper surveys the state-of-the-art techniques and models developed for text-to-video generation, exploring approaches that convert semantic information from written text into visual sequences. We review various methods, including generative adversarial networks (GANs), transformers, and diffusion models, discuss their architectures, training strategies, and challenges in generating coherent and temporally consistent video. Finally, the survey highlights current limitations, including dataset, DL model, evaluation parameter, scalability, and computational requirements.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Survey on Text-to-Video Generation: Models, Architectures, and Challenges

  • Ravikumar Patel,
  • Hemit Rana,
  • Nikita Bhatt

摘要

The generation of video from textual descriptions is an emerging interdisciplinary research area, merging advancements in natural language processing (NLP), computer vision, and deep learning. This paper surveys the state-of-the-art techniques and models developed for text-to-video generation, exploring approaches that convert semantic information from written text into visual sequences. We review various methods, including generative adversarial networks (GANs), transformers, and diffusion models, discuss their architectures, training strategies, and challenges in generating coherent and temporally consistent video. Finally, the survey highlights current limitations, including dataset, DL model, evaluation parameter, scalability, and computational requirements.