DPG: Deterministic Policy Gradient
摘要
When calculating policy gradient using the vanilla PG algorithm, we need to take the expectation over states as well as actions.