An Actor Critic Machine Learning Method for Free Terminal Time Optimal Control
摘要
Optimal feedback control of nonlinear system with free terminal time present many challenges including nonsmooth in the value function and control laws, and existence of multiple local or even global optimal trajectories. To mitigate these issues, the authors introduce an actor-critic method along with some enhancements. The authors demonstrate the algorithm’s effectiveness on a prototypical example featuring each of the main pathological issues present in problems of this type as well as a higher dimensional example to show that the solution method presented can scale.