One of the most urgent challenges in Reinforcement Learning research is the lack of reproducibility. Therefore, to further the understanding of the training behavior of Reinforcement Learning agents, we analyze the training of agents playing the established baseline environment Taxi. In particular, we contrast results based on different forms of exploration. In addition, we can demonstrate that in this context penalization without termination is to be the preferred punishment for incorrect actions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Statistical Analysis of Reinforcement Learning Training

  • Maximilian Moll,
  • Matthias Schilling,
  • Stefan Pickl

摘要

One of the most urgent challenges in Reinforcement Learning research is the lack of reproducibility. Therefore, to further the understanding of the training behavior of Reinforcement Learning agents, we analyze the training of agents playing the established baseline environment Taxi. In particular, we contrast results based on different forms of exploration. In addition, we can demonstrate that in this context penalization without termination is to be the preferred punishment for incorrect actions.