MCRE: Multimodal Conditional Representation and Editing for Text-Motion Generation
摘要
Recent advancements in text-to-motion generation models have shown impressive capabilities in creating high-fidelity motion sequences. However, generating desired sequences using only text prompts is challenging due to the complexity of prompt engineering. We present the Multimodal Conditional Representation and Editing (MCRE) module, a lightweight adapter for text-to-motion generation and editing. MCRE unifies text and motion conditions into the CLIP representation space, enabling precise and flexible multimodal control. Despite its simplicity, MCRE’s hybrid motion and text-conditioned editing capabilities achieve comparable or better performance than fully fine-tuned models. The highly disentangled CLIP representations enable flexible motion sequence editing by combining multiple conditions, resulting in versatile and high-quality motion generation and editing.