This paper introduces a method to predict pedestrian trajectories based on their policies. Current airport pedestrian trajectory prediction methods have challenges: supervised learning predictions work only for the same scene and fail when the scene changes; reinforcement learning requires setting a reasonable reward function; without trajectory data, inverse reinforcement learning cannot directly learn the reward function. To solve the first challenge, the paper considers learning pedestrian policies from the original environment. Then, it applies these policies in different environments for trajectory prediction. To address the last two challenges, the paper proposes the Generative Adversarial Target Alignment (GATA) framework. This framework learns pedestrian policies and reward functions simultaneously during target alignment. Experiments show that GATA effectively achieves target alignment, pedestrian policy learning, and reward function learning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generation of Airport Pedestrian Boarding Decision Process Based on Target Alignment

  • Weifeng Xu,
  • Qian Luo,
  • Wanli Dang,
  • Long Gen,
  • Jiaoling Zheng

摘要

This paper introduces a method to predict pedestrian trajectories based on their policies. Current airport pedestrian trajectory prediction methods have challenges: supervised learning predictions work only for the same scene and fail when the scene changes; reinforcement learning requires setting a reasonable reward function; without trajectory data, inverse reinforcement learning cannot directly learn the reward function. To solve the first challenge, the paper considers learning pedestrian policies from the original environment. Then, it applies these policies in different environments for trajectory prediction. To address the last two challenges, the paper proposes the Generative Adversarial Target Alignment (GATA) framework. This framework learns pedestrian policies and reward functions simultaneously during target alignment. Experiments show that GATA effectively achieves target alignment, pedestrian policy learning, and reward function learning.