Learning Social and Physical Compliant Multi-modal Futures
摘要
Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning. It is not easy for prediction systems to successfully model the inherent uncertainty of futures while take into account dynamic social and physical interactions in a highly interactive circumstance. Towards these goals, we develop a unifying model to predict socially and physically acceptable multi-modes of future trajectories. The modes of trajectories condition on past observation, which are not explicit labels but indicate walking patterns (such as walking straight, turning left/right) and interacting patterns (such as aggressive, mild). Our model contains an encoder and a decoder, which leverages two source of information, all past path of agents in a shared scenario and scene context to jointly learn the representations of social and physical interactions, where the former can efficiently scale to any number of agents and the latter helps the model learn the traversable parts of the scenario. Extensive experiments over several trajectory prediction benchmarks demonstrate that our method is able to capture the multi-modality of human motion and forecast the distributions of plausible futures in complex scenarios.