<p>Deep Reinforcement learning (DRL) is used to enable autonomous navigation in unknown environments. Most research assumes perfect sensor data, but real-world environments may contain natural and artificial sensor noise and denial. Here, we present a benchmark of both well-used and emerging DRL algorithms in two navigation tasks - Lidar + position, and vision end-to-end - with configurable sensor denial effects. In particular, we are interested in comparing how different DRL methods (e.g. model-free, on-policy PPO vs. model-free off-policy TD3, vs. model-based DreamerV3) are affected by imperfect sensor readings. We show that DreamerV3 outperforms other methods in the visual end-to-end navigation task with a dynamic goal. Furthermore, DreamerV3 generally outperforms other methods in sensor-denied environments. In order to improve robustness, we use adversarial training and demonstrate an improved performance in denied environments, although we show that this may lead to the agent learning to choose high-risk actions in case of uncertain sensor readings, which is not appropriate for safety-critical scenarios. We anticipate this benchmark of different DRL methods and the usage of adversarial training to be a starting point for the development of more elaborate navigation strategies that are capable of dealing with uncertain and denied sensor readings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Deep Reinforcement Learning for Navigation in Denied Sensor Environments

  • Mariusz Wisniewski,
  • Paraskevas Chatzithanos,
  • Weisi Guo,
  • Antonios Tsourdos

摘要

Deep Reinforcement learning (DRL) is used to enable autonomous navigation in unknown environments. Most research assumes perfect sensor data, but real-world environments may contain natural and artificial sensor noise and denial. Here, we present a benchmark of both well-used and emerging DRL algorithms in two navigation tasks - Lidar + position, and vision end-to-end - with configurable sensor denial effects. In particular, we are interested in comparing how different DRL methods (e.g. model-free, on-policy PPO vs. model-free off-policy TD3, vs. model-based DreamerV3) are affected by imperfect sensor readings. We show that DreamerV3 outperforms other methods in the visual end-to-end navigation task with a dynamic goal. Furthermore, DreamerV3 generally outperforms other methods in sensor-denied environments. In order to improve robustness, we use adversarial training and demonstrate an improved performance in denied environments, although we show that this may lead to the agent learning to choose high-risk actions in case of uncertain sensor readings, which is not appropriate for safety-critical scenarios. We anticipate this benchmark of different DRL methods and the usage of adversarial training to be a starting point for the development of more elaborate navigation strategies that are capable of dealing with uncertain and denied sensor readings.