Modulating chaos in spatiotemporal systems based on deep reinforcement learning
摘要
In this paper, the deep reinforcement learning technique is applied to modulating chaos in spatiotemporal systems without having any prior knowledge of their underlying dynamics. The deep deterministic policy gradient includes two modules, Actor and Critic, which are composed of four artificial neural networks to improve learning stability and output deterministic actions. This framework successfully learns to suppress or activate spatiotemporal chaotic patterns, relying on nothing but a reward for asynchronous states. Indeed, the reward function is the key to the successful application of deep reinforcement learning. We design a reward function based on the moving largest Lyapunov exponents relying on a fixed-length set consisting of recent states evolving over time, which only involves a small subset of the system’s variables. Considering the Kuramoto–Sivashinsky system as the test model, we show this type of reward function is proved to be useful and universal both for suppression and for activation of chaos.