Policy-Based Gradient-Free Algorithms
摘要
So far, we have used the parametric function 𝜋(𝛉) to approximate optimal policy, and try to find suitable policy parameter 𝛉 to maximize the return expectation.