<p>Reinforcement learning (RL)-based control in tokamaks offers improved flexibility for nuclear fusion, but typically depends on simulators that can accurately evolve the high-dimensional plasma state. First-principle simulators are often too computationally intensive for efficient RL training. Here, we develop a fully data-driven simulator that mitigates compounding errors caused by its autoregressive nature. This high-fidelity model enables rapid training of an RL agent that generates engineering-reasonable actuator commands to reach long-term plasma configuration targets. Combined with a neural network surrogate for equilibrium fitting (EFITNN), the agent maintains a 400-ms, 1 kHz control trajectory on the HL-3 tokamak, accurately tracking plasma current and boundary shape. It also adapts to changes in triangularity without retraining, demonstrating robustness. These results show that data-driven dynamics models can support fast and reliable RL-based control, meeting anticipated engineering requirements for routine operation in future fusion devices such as ITER.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-fidelity data-driven dynamics model for reinforcement learning-based control in HL-3 tokamak

  • Niannian Wu,
  • Zongyu Yang,
  • Rongpeng Li,
  • Ning Wei,
  • Yihang Chen,
  • Qianyun Dong,
  • Jiyuan Li,
  • Guohui Zheng,
  • Xinwen Gong,
  • Feng Gao,
  • Bo Li,
  • Min Xu,
  • Zhifeng Zhao,
  • Wulyu Zhong

摘要

Reinforcement learning (RL)-based control in tokamaks offers improved flexibility for nuclear fusion, but typically depends on simulators that can accurately evolve the high-dimensional plasma state. First-principle simulators are often too computationally intensive for efficient RL training. Here, we develop a fully data-driven simulator that mitigates compounding errors caused by its autoregressive nature. This high-fidelity model enables rapid training of an RL agent that generates engineering-reasonable actuator commands to reach long-term plasma configuration targets. Combined with a neural network surrogate for equilibrium fitting (EFITNN), the agent maintains a 400-ms, 1 kHz control trajectory on the HL-3 tokamak, accurately tracking plasma current and boundary shape. It also adapts to changes in triangularity without retraining, demonstrating robustness. These results show that data-driven dynamics models can support fast and reliable RL-based control, meeting anticipated engineering requirements for routine operation in future fusion devices such as ITER.