Conditional Variational Inference for Multi-modal Trajectory Prediction with Latent Diffusion Prior
摘要
Predicting pedestrian trajectories is vital for improving safety and efficiency in human-robot interaction within traffic systems. However, this task is inherently challenging due to the unpredictable nature of human behavior. We present MotDiff, a method based on Variational Auto-encoders with a diffusion prior, which synthesizes latent variables to capture the unobserved uncertainty and complex relation among agents. We provide a comprehensive theoretical background of our approach and evaluate it with various generative modeling methods using three public pedestrian datasets, showing its effectiveness in achieving both accuracy and diversity.