Towards Efficient Reinforcement Learning: A Transfer Approach via Task and Environment Difference Elimination
摘要
Enhancing learning efficiency and convergence performance through the transfer of external knowledge is a common paradigm in reinforcement learning. Previous research in cross-domain knowledge transfer often discusses the task or environment differences between the source and target domains in isolation. However, the two types of differences typically coexist. ‘In this paper, we introduce a novel transfer reinforcement learning method, Task and Environment Difference Elimination (TEDE), aimed at improving the learning efficiency by concurrently mitigating task and environment differences. Specifically, TEDE captures environment and task differences through marginal and conditional distribution discrepancies, thereby simultaneously eliminating the differences arising from varying tasks and environments. This capability enables TEDE to facilitate the transfer of cross-domain generalizable knowledge. We substantiate the efficacy of our proposed method in the Google Research Football environment. Experiments demonstrate that TEDE is capable of extracting deep, high-level features that possess both environment generalizability and task adaptability, leading to a significant enhancement in the learning efficiency and convergence performance of the agents.