<p>This paper asks the question of whether there are performance reasons for leveraging Human Values in the pre-training of Modern Reinforcement Learning (RL) models? As part of exploring this question, this paper looks at Modern RL algorithms generally, and specifically Model Based and Model Free RL. It then examines the treatment of Rewards in such model types as the Solution Concepts of: Common Reward Games, Zero Sum Games, and General Sum Games (incl. Nash Equilibriums). The Value Alignment Problem is then described as a result of a Zero Sum Game Solution Concepts and actual examples of this problem are provided. The paper then goes on to explore how Ruchard Sutton and Stewart Russell propose to address the Value Alignment Problem. Finally, the paper examines the possible use of Supervised Learning to effectively pre-train Modern RL algorithms to address the Value Alignment Problem, and cites the success of the AlphaStar algorithm as an example how pre-training with Human Values may have technical benefits.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond good and evil?: understanding the role of human values in modern reinforcement learning

  • Theodore McCullough

摘要

This paper asks the question of whether there are performance reasons for leveraging Human Values in the pre-training of Modern Reinforcement Learning (RL) models? As part of exploring this question, this paper looks at Modern RL algorithms generally, and specifically Model Based and Model Free RL. It then examines the treatment of Rewards in such model types as the Solution Concepts of: Common Reward Games, Zero Sum Games, and General Sum Games (incl. Nash Equilibriums). The Value Alignment Problem is then described as a result of a Zero Sum Game Solution Concepts and actual examples of this problem are provided. The paper then goes on to explore how Ruchard Sutton and Stewart Russell propose to address the Value Alignment Problem. Finally, the paper examines the possible use of Supervised Learning to effectively pre-train Modern RL algorithms to address the Value Alignment Problem, and cites the success of the AlphaStar algorithm as an example how pre-training with Human Values may have technical benefits.