Optimal control under safety constraints and disturbances: a multi-step, off-policy adaptive dynamic programming approach
摘要
This paper introduces a multi-step, off-policy adaptive dynamic programming approach, in both model-free and model-based variants, intending to solve optimal control problems under disturbances and safety constraints. To provide a more accurate estimation of the performance function in the policy evaluation step, we employ an interleaved training method in the model-free scheme and utilize a prior model in the model-based version to mitigate the underestimation issue of the accumulated utility function. To further counteract the underestimation of the terminal performance function, dual critic neural networks are utilized. Additionally, to ensure a well-balanced trade-off between safety and performance requirements, the original unconstrained policy improvement process is transformed into a constrained optimization task with a far-sighted safety function. Furthermore, an actor-critic-disturbance framework is designed to handle safety constraints during the zero-sum game process, in which the disturbance policy and the performance function are alternately updated during the PEV step. Based on this, a rigorous theoretical analysis is conducted to evaluate the convergence property of the proposed method. Finally, simulation results and practical experiments demonstrate the effectiveness and safety of the proposed method.