Recent advancements in text-to-motion generation models have shown impressive capabilities in creating high-fidelity motion sequences. However, generating desired sequences using only text prompts is challenging due to the complexity of prompt engineering. We present the Multimodal Conditional Representation and Editing (MCRE) module, a lightweight adapter for text-to-motion generation and editing. MCRE unifies text and motion conditions into the CLIP representation space, enabling precise and flexible multimodal control. Despite its simplicity, MCRE’s hybrid motion and text-conditioned editing capabilities achieve comparable or better performance than fully fine-tuned models. The highly disentangled CLIP representations enable flexible motion sequence editing by combining multiple conditions, resulting in versatile and high-quality motion generation and editing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCRE: Multimodal Conditional Representation and Editing for Text-Motion Generation

  • Tengjiao Sun,
  • Xiang Li,
  • Tianyu Shi,
  • Jiahui Peng,
  • Sheng Zheng,
  • Hansung Kim

摘要

Recent advancements in text-to-motion generation models have shown impressive capabilities in creating high-fidelity motion sequences. However, generating desired sequences using only text prompts is challenging due to the complexity of prompt engineering. We present the Multimodal Conditional Representation and Editing (MCRE) module, a lightweight adapter for text-to-motion generation and editing. MCRE unifies text and motion conditions into the CLIP representation space, enabling precise and flexible multimodal control. Despite its simplicity, MCRE’s hybrid motion and text-conditioned editing capabilities achieve comparable or better performance than fully fine-tuned models. The highly disentangled CLIP representations enable flexible motion sequence editing by combining multiple conditions, resulting in versatile and high-quality motion generation and editing.