Closed-Loop Safe Correction for Reinforcement Learning Policy
摘要
Trial and error learning is an approach with uncertain consequences. How to maintain policy security, stability, and efficiency under controlled circumstances, posing a significant academic challenge. Such as Reinforcement Learning (RL) leverages the neural network parameter space to identify the optimal solution for achieving the optimal mapping of the policy space. However, even a slight fluctuation in the parameter space can lead to policy collapse. To solve it, we propose a method called Closed-loop Safe Correction (CLSC). This method employs twin agents to establish a closed-loop framework for policy and sampling data distribution, gradually improving policy efficiency while ensuring policy safety. Base on theoretical and experimental analysis, CLSC elucidates the performance, influencing factors, and advantages of this approach in policy safety scenarios. These findings provide novel research insights and practical directions for advancing safe RL.