Deep Reinforcement Learning for Thermal Soaring of Small Gliders
摘要
Compared with large gliders, small gliders experience a more substantial reduction in elevator effectiveness while maneuvering to locate and exploit thermals, primarily due to their lower Reynolds numbers, smaller horizontal tail, and higher sensitivity of flow attachment. In this study, deep reinforcement learning is employed to train a small glider with a wingspan of 3.4 m to explore and exploit thermals. The agent interacts with a high-fidelity nonlinear flight dynamics environment, with its control policy implemented via a Long Short-Term Memory (LSTM) network and optimized using Proximal Policy Optimization (PPO). The flight trajectories in the tests exhibit characteristic spiral ascents, demonstrating successful thermal soaring in sparse-reward scenarios. The results highlight the robustness and applicability of the approach, while the trained agent can be deployed on a companion computer and further fine-tuned through interaction with real-world environments.