错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DPG: Deterministic Policy Gradient

  • Zhiqing Xiao

摘要

When calculating policy gradient using the vanilla PG algorithm, we need to take the expectation over states as well as actions.