Action Conditioned Attention Encoder-Decoder and Discriminator for Human Motion Generation
摘要
We present a CVAE-GAN-based architecture for human motion generation with an action-conditioned variational autoencoder and a generative discriminator. In this work, we focus on generating more accurate actions performed by a single person using a conditioned generative model. The primary motivation for this work comes from human-robot collaboration scenarios where a person interacts with the robot using human actions. Our approach consists of a self-attention-based conditional variational autoencoder for reconstructions and a graph network-based discriminator for realistic human motion quality. We evaluate our network on three open-source datasets known as HumanAct12, NTU 120 RGB+D, and Actions in Supermarket Dataset. The extensive experiments show that the presented approach works for various human motions and input representations, such as the SMPL pose parameters, trajectory data, and skeleton joints. We achieve higher accuracy compared to the state-of-the-art methods on all three datasets.