Deep reinforcement learning is a highly suitable method for real-time control under uncertain conditions based on the Markov Decision Process (MDP), but its application to the stochastic dynamic optimal power flow (SDOPF) program still remains challenging due to its limitations in meeting constraints under uncertainty. While pioneering researched have explored constrained MDP and risk-aware MDP formulations for SDOPF, aiming to minimize cumulative constraint violations, they all face difficulties in satisfying state-wise safety constraints. This paper proposes a chance constrained MDP formulation for SDOPF and a Bayesian primal-dual safe policy optimization approach. Using Bayesian estimation to update the posterior distribution of both state- and trajectory-wise constraint violations, and combining them with the advantage function, the efficiency and safety of the policy are enhanced. Case studies validate the cost-effectiveness and comprehensive safety performance of the proposed formulation and solution compared to state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bayesian Primal-Dual Safe Policy Optimization for Stochastic Dynamic Optimal Power Flow Program: Dissecting from a Probabilistic Perspective

  • Yujian Ye,
  • Yizhi Wu,
  • Jianxiong Hu,
  • Qiong Wang,
  • Xi Zhang,
  • Goran Strbac

摘要

Deep reinforcement learning is a highly suitable method for real-time control under uncertain conditions based on the Markov Decision Process (MDP), but its application to the stochastic dynamic optimal power flow (SDOPF) program still remains challenging due to its limitations in meeting constraints under uncertainty. While pioneering researched have explored constrained MDP and risk-aware MDP formulations for SDOPF, aiming to minimize cumulative constraint violations, they all face difficulties in satisfying state-wise safety constraints. This paper proposes a chance constrained MDP formulation for SDOPF and a Bayesian primal-dual safe policy optimization approach. Using Bayesian estimation to update the posterior distribution of both state- and trajectory-wise constraint violations, and combining them with the advantage function, the efficiency and safety of the policy are enhanced. Case studies validate the cost-effectiveness and comprehensive safety performance of the proposed formulation and solution compared to state-of-the-art approaches.