Inverse optimal control of general nonlinear systems based on adaptive dynamic programming
摘要
This paper proposes a novel model-free inverse optimal control (IOC) method for general nonlinear continuous-time systems based on Q-learning, which overcomes the limitations of existing approaches that rely on affine system structures and explicit model information. First, an inner-outer two-layer iterative framework is designed, where the inner layer approximates the optimal control policy and Q-function, while the outer layer identifies the cost function parameters from expert trajectories. Then, for the first time, a rigorous theoretical convergence analysis of the Q-learning algorithm for continuous-time nonlinear systems is provided. And the policy reconstruction and cost function inference are achieved using only expert data. The paper also employs neural networks to construct an actor-critic architecture and updates weights via gradient-based optimization. Finally, numerical simulations are conducted to verify the effectiveness of the proposed algorithm, offering new theoretical foundations and technical pathways for inverse optimal control of nonlinear systems.