Adaptive Self-supervised Agent: Leveraging Deep Reinforcement Learning Through Strategic Precedence Estimation
摘要
In this paper, we explore the idea of fueling the agent policy with accurate trajectories instead of learning these trajectories through interaction with the environment. This mechanism is entirely based on the precedence estimation that gives the agent policy the flexibility to learn the best experience. In other words, instead of starting from scratch with no experience of the environment, we inject accurate trajectories to start the policy learning and tuning. This procedure has progressively and steadily enhanced the action quality. It pushes the policy towards actions that deliver a high execution. Extensive results are given to confirm the efficiency of the proposed approach. This basic idea has achieved a promising performance in comparison with sophisticated approaches.