Deep reinforcement learning (DRL) still explores insufficiently when dealing complex tasks with high-dimensional and large state spaces. Therefore, developing better exploration strategies is still one of the important tasks in reinforcement learning. The paper introduces a new exploration strategy ATLF (Adjustment of Truncated Lévy Flight exploration framework, ATLF) which augments the existing exploration mode with Lévy flight, making action selection more stochastic to boost exploration. The ALTF framework is combined with discrete-space algorithm DQN and continuous space-algorithm SAC to handle reinforcement learning tasks. Compared with a variety of reinforcement learning algorithms on OpenAI gym environments such as MountainCar-v0 and Walker2d-v2, the result shows that our algorithm has better exploration ability than vanilla DQN or SAC, obtaining higher overall rewards, and is less likely to fall into local optimization, and is more stable. Additionally, the result shows that the ALTF is highly compatible with existing deep reinforcement learning algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Framework of Reinforcement Learning for Truncated Lévy Flight Exploratory

  • Quan Liu,
  • Shile Feng,
  • Zixian Gu

摘要

Deep reinforcement learning (DRL) still explores insufficiently when dealing complex tasks with high-dimensional and large state spaces. Therefore, developing better exploration strategies is still one of the important tasks in reinforcement learning. The paper introduces a new exploration strategy ATLF (Adjustment of Truncated Lévy Flight exploration framework, ATLF) which augments the existing exploration mode with Lévy flight, making action selection more stochastic to boost exploration. The ALTF framework is combined with discrete-space algorithm DQN and continuous space-algorithm SAC to handle reinforcement learning tasks. Compared with a variety of reinforcement learning algorithms on OpenAI gym environments such as MountainCar-v0 and Walker2d-v2, the result shows that our algorithm has better exploration ability than vanilla DQN or SAC, obtaining higher overall rewards, and is less likely to fall into local optimization, and is more stable. Additionally, the result shows that the ALTF is highly compatible with existing deep reinforcement learning algorithms.