This paper explores the under-examined potential of offline reinforcement learning algorithms in the context of Smart Grids. While online methods, such as Proximal Policy Optimization (PPO), have been extensively studied, offline methods, which inherently avoid real-time interactions, may offer practical safety benefits in scenarios like power grid management, where suboptimal policies could lead to severe consequences. To investigate this, we conducted experiments in Grid2Op environments with varying grid complexity, including differences in size and topology. Our results suggest that offline algorithms can achieve comparable or superior performance to online methods, particularly as grid complexity increases. Additionally, we observed that the diversity of training data plays a crucial role, with data collected through environment sampling yielding better results than data generated by trained models. These findings underscore the value of further exploring offline approaches in safety-critical applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Offline Reinforcement Learning Approaches for Safe and Effective Smart Grid Control

  • Angel Peredo,
  • Hector Lugo,
  • Christian Narcia-Macias,
  • Jose Espinoza,
  • Daniel Masamba,
  • Adan Gandarilla,
  • Erik Enriquez,
  • Dong-Chul Kim

摘要

This paper explores the under-examined potential of offline reinforcement learning algorithms in the context of Smart Grids. While online methods, such as Proximal Policy Optimization (PPO), have been extensively studied, offline methods, which inherently avoid real-time interactions, may offer practical safety benefits in scenarios like power grid management, where suboptimal policies could lead to severe consequences. To investigate this, we conducted experiments in Grid2Op environments with varying grid complexity, including differences in size and topology. Our results suggest that offline algorithms can achieve comparable or superior performance to online methods, particularly as grid complexity increases. Additionally, we observed that the diversity of training data plays a crucial role, with data collected through environment sampling yielding better results than data generated by trained models. These findings underscore the value of further exploring offline approaches in safety-critical applications.