A motion conditioned diffusion model for real-time hand trajectory semantic prediction
摘要
Vision-based real-time gesture trajectory semantic prediction is challenging due to inherent ambiguity on semantic relatedness, which often result in high uncertainty in the dynamic and continuous semantic recognition scenes. In this paper, a motion-conditioned diffusion model for skeleton-based hand trajectory semantic prediction (DiffHand) is proposed. First, to improve the computational efficiency, we input the coordinates of the skeletal points representing the hand pose into the model. Then, we add random noise to the real future sequences to form a random noise sequence that conforms to a normal distribution. Next, we encode and feature fuse the past sequences. Finally, we use conditional features and random noise sequences to guide the model in generating continuous and reliable predicted motion trajectories. Experimental results show that our method has a high prediction accuracy of 76.3% and a recognition speed of 33 fps. Our method strikes a good balance between prediction accuracy and recognition speed.