Reinforcement Learning for Short-Term Battery Dispatch in Renewable-Thermal Power Systems
摘要
This paper presents a Reinforcement Learning (RL) framework for short-term battery dispatch in renewable-thermal power systems under uncertainty. The problem is formulated as a Markov decision process over a 14-day hourly horizon, where an agent selects charging and discharging actions to minimize thermal generation costs. A custom Gymnasium-compatible environment was implemented, and tabular Q-learning with \(\epsilon \) -greedy exploration was trained on 115 stochastic residual-demand scenarios calibrated to the Uruguayan grid. The learned policies converged reliably and displayed interpretable behavior, charging during low-cost hours and discharging at peaks to reduce reliance on expensive thermal units. Compared with benchmark strategies, the RL agent reduced thermal generation costs by 2.5% relative to the no-battery case and remained within 1.3% of an optimization-based reference, requiring under five hours on a standard workstation. The formulation is simple, reproducible, and extensible, positioning RL as a lightweight complement to classical optimization methods for power system operation under uncertainty.