<p>Robot crowd navigation has been gaining increasing attention and popularity in various practical applications. In the existing research, deep reinforcement learning has been applied to robot crowd navigation by training policies in an online mode. However, this inevitably leads to unsafe exploration and consequently causes low sampling efficiency during pedestrian–robot interaction. To this end, this paper proposes an offline spatial–temporal actor-critic algorithm for robot crowd navigation by utilizing pre-collected crowd navigation experience. Specifically, a novel SARSA-like loss function is defined to update the separate value function, eliminating the need for behavior regularization or constraints. Furthermore, this algorithm incorporates a spatial–temporal transformer into offline actor-critic learning that captures the spatial–temporal features from the offline pedestrian–robot interactions. It allows our robot navigation policy to exhibit greater adaptability and less conservatism to the highly dynamic crowd environments. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods by means of qualitative and quantitative analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Offline spatial–temporal actor-critic learning for robot crowd navigation without behavior regularization

  • Shuai Zhou,
  • Hao Fu,
  • Haodong He,
  • Wei Liu,
  • Zixin Huang

摘要

Robot crowd navigation has been gaining increasing attention and popularity in various practical applications. In the existing research, deep reinforcement learning has been applied to robot crowd navigation by training policies in an online mode. However, this inevitably leads to unsafe exploration and consequently causes low sampling efficiency during pedestrian–robot interaction. To this end, this paper proposes an offline spatial–temporal actor-critic algorithm for robot crowd navigation by utilizing pre-collected crowd navigation experience. Specifically, a novel SARSA-like loss function is defined to update the separate value function, eliminating the need for behavior regularization or constraints. Furthermore, this algorithm incorporates a spatial–temporal transformer into offline actor-critic learning that captures the spatial–temporal features from the offline pedestrian–robot interactions. It allows our robot navigation policy to exhibit greater adaptability and less conservatism to the highly dynamic crowd environments. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods by means of qualitative and quantitative analysis.