Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
摘要
This extended abstract summarizes [10], and cites it as the primary source. We present a quantum-compatible reformulation of Natural Policy Gradient (NPG) using deterministic, truncation-based estimators for both the policy gradient and Fisher information, enabling coherent evaluation with standard quantum environment and policy oracles. Leveraging quantum mean estimation and variance reduction within a classical–quantum double-loop scheme, the resulting Quantum NPG (QNPG) achieves a sample complexity of \(\tilde{\mathcal {O}}(\epsilon ^{-1.5})\) to reach an \(\epsilon \) -optimal policy, improving upon the classical \({\varOmega }(\epsilon ^{-2})\) lower bound for policy-gradient methods. We outline the required assumptions (smooth score function, Fisher non-degeneracy, bounded compatible approximation error), highlight the exponentially decaying truncation bias, and sketch the convergence analysis that yields the overall complexity. This manuscript is intended as a concise, approximately six-page overview. We refer the full details, proofs, and extended discussion to [10].