Efficient Hierarchical Reinforcement Learning via Mutual Information Constrained Subgoal Discovery
摘要
Goal-conditioned hierarchical reinforcement learning has demonstrated impressive capabilities in addressing complex and long-horizon tasks. However, the extensive subgoal space often results in low sample efficiency and challenging exploration. To address this issue, we extract informative subgoals by constraining their generation range in mutual information distance space. Specifically, we impose two constraints on the high-level policy during off-policy training: the generated subgoals should be reached with less effort by the low-level policy, and the realization of these subgoals can facilitate achieving the desired goals. These two constraints enable subgoals to act as critical links between the current states and the desired goals, providing more effective guidance to the low-level policy. The empirical results on continuous control tasks demonstrate that our proposed method significantly enhances the training efficiency, regardless of the dimensions of the state and action spaces, while ensuring comparable performance to state-of-the-art methods.