Implicit Cooperative Trajectory Planning with Learned Rewards Under Uncertainty
摘要
Urban traffic scenarios often require high interaction between traffic participants to ensure safety and efficiency. While the capabilities of automated driving systems have made remarkable progress in the past decade, they lack two critical abilities: anticipation and provision of cooperation between traffic participants without communication, i.e., implicit cooperation. Observing the behavior of other traffic participants, humans infer the need to cooperate and act accordingly. Our work presents a system that utilizes a sampling-based cooperative trajectory planner that accounts for all possible actions of other traffic participants, enabling cooperation. Further, we extend the planner employing learned reward models based on expert trajectories to demonstrate its ability to adapt to a desired human driving style for smooth integration into today’s traffic. Lastly, we address the issue of measurement uncertainties to robustify the decision-making process in real-world environments utilizing return distributions over start states according to a belief. We exemplify the effectiveness of our solutions on 15 challenging multi-agent scenarios.