Bayesian Primal-Dual Safe Policy Optimization for Stochastic Dynamic Optimal Power Flow Program: Dissecting from a Probabilistic Perspective
摘要
Deep reinforcement learning is a highly suitable method for real-time control under uncertain conditions based on the Markov Decision Process (MDP), but its application to the stochastic dynamic optimal power flow (SDOPF) program still remains challenging due to its limitations in meeting constraints under uncertainty. While pioneering researched have explored constrained MDP and risk-aware MDP formulations for SDOPF, aiming to minimize cumulative constraint violations, they all face difficulties in satisfying state-wise safety constraints. This paper proposes a chance constrained MDP formulation for SDOPF and a Bayesian primal-dual safe policy optimization approach. Using Bayesian estimation to update the posterior distribution of both state- and trajectory-wise constraint violations, and combining them with the advantage function, the efficiency and safety of the policy are enhanced. Case studies validate the cost-effectiveness and comprehensive safety performance of the proposed formulation and solution compared to state-of-the-art approaches.