Discarding Erroneous Knowledge Online in Transfer Reinforcement Learning
摘要
This study deals with transfer reinforcement learning, i.e., reinforcement learning using knowledge acquired in one learning process (source task) in other learning processes (target tasks). The knowledge improves performance in the early stages of learning, but it may sometimes be detrimental, which is called negative transfer. To avoid the negative transfer, we propose a method in which the agent itself measures an effect of the transferred knowledge on the current learning process, and if it is harmful, it discards the knowledge during learning. We conducted experiments in a multi-player game environment with independently learning agents, in which the agents first learned their policies in one map of the environment and after that they learned in other, randomly generated maps. The result shows that the proposal mitigated negative transfer more successfully than existing methods.