错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Q-Learning from Statistical Demonstrations for Efficient Battery Storage Energy Arbitrage

  • Chois Zhi Cai,
  • Hui Tian,
  • Leo Yu Zhang,
  • Lee Weng

摘要

This paper presents a novel Reinforcement Learning (RL) approach—Deep Q-learning from Statistical Demonstrations (DQfSD)—for Battery Energy Storage Systems (BESSs) to enable effective energy arbitrage operations. The method uses minimal human expertise and a statistically derived rule-based policy as demonstration data. The policy guides early learning and improves sample efficiency compared to standard RL methods. We evaluate DQfSD against three widely used RL algorithms: Deep Q-Network, Proximal Policy Optimization, and Asynchronous Advantage Actor-Critic, using the real-world wholesale price data from AEMO (Australian Energy Market Operator) Queensland. The simulated BESS compromises a 2 MWh battery with a 1 MW inverter, operated over a 12-month test period. In the most challenging benchmark, DQfSD was trained on only three months of data, while the baselines were trained on six months of data with up to 1000 episodes. DQfSD achieved a total reward outperforming the best baseline by 11.2%. Operational metrics confirm that these gains were achieved without increasing battery cycling. The complete implementation, including the simulation environment, data processing scripts, and experiment configurations, is released as open source to support reproducibility and further research.