IT-HMDM: Invertible Transformer for Human Motion Diffusion Model
摘要
Generating realistic and natural human motions has been a challenging task. Despite decades of research on modeling human motions, synthesizing realistic and natural sequences remains extremely challenging. In this paper, we propose an invertible Transformer for human motions with diffusion model (IT-HMDM). The model takes into account the input text lexicality and uses lexical encoding to enhance the feature capturing in hidden space. It also uses bijective affine transformation and logarithmic determinant regular terms to reduce information loss during the encoding process. We also use an improved grouping attention mechanism for semantic injection in order to reduce computation complexity and provide better model performance. The experimental results prove the validation of our model. When tested on the HumanML3D datasets, our model improves over MDM in R Precision (top 3), FID, and Multimodal Dist by 5%, 27.6%, and 6%, respectively.