错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Policy-Based Gradient-Free Algorithms

  • Zhiqing Xiao

摘要

So far, we have used the parametric function 𝜋(𝛉) to approximate optimal policy, and try to find suitable policy parameter 𝛉 to maximize the return expectation.