High-fidelity data-driven dynamics model for reinforcement learning-based control in HL-3 tokamak
摘要
Reinforcement learning (RL)-based control in tokamaks offers improved flexibility for nuclear fusion, but typically depends on simulators that can accurately evolve the high-dimensional plasma state. First-principle simulators are often too computationally intensive for efficient RL training. Here, we develop a fully data-driven simulator that mitigates compounding errors caused by its autoregressive nature. This high-fidelity model enables rapid training of an RL agent that generates engineering-reasonable actuator commands to reach long-term plasma configuration targets. Combined with a neural network surrogate for equilibrium fitting (EFITNN), the agent maintains a 400-ms, 1 kHz control trajectory on the HL-3 tokamak, accurately tracking plasma current and boundary shape. It also adapts to changes in triangularity without retraining, demonstrating robustness. These results show that data-driven dynamics models can support fast and reliable RL-based control, meeting anticipated engineering requirements for routine operation in future fusion devices such as ITER.