Semantic-Driven Free-View 3D Human Motion Video Composite
摘要
We propose a text-to-video synthesis method with controllable free perspective. This method performs well for human-centered textual descriptions and ensures the stability and realism of the person's motion from any perspective in the generated videos. Specifically, we analyze the input textual description and decompose it into three retrieval instructions. Different retrieval methods are applied based on the type of retrieval, targeting the key elements from our constructed resource library, including people, backgrounds, and motion sequences. These retrieval methods ensure semantic consistency in video synthesis. Furthermore, we construct a foreground character library using multi-view RGB images and leverage the advantages of 3D reconstruction to implicitly model the retrieved foreground characters. This ensures the stability of the characters during video synthesis and enables free-perspective transformations. To address the limitation of existing methods in generating complex motions, we employ real motion sequences to drive the reconstruction, achieving video synthesis of arbitrary duration. Experimental results demonstrate that our method outperforms existing open-source models across multiple metrics.