<p>We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where the controls applied to the diffusion term satisfy a smallness condition. We demonstrate the convergence of PIA based on a uniform <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="245_2025_10249_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="30" /> </InlineMediaObject> <EquationSource Format="TEX">\({{\mathcal {C}}}^{2,\alpha }\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mrow> <mi mathvariant="script">C</mi> </mrow> <mrow> <mn>2</mn> <mo>,</mo> <mi>α</mi> </mrow> </msup> </math></EquationSource> </InlineEquation> estimate for the value sequence generated by PIA, and provide a quantitative convergence analysis for this scenario. Second, we investigate PIA with unbounded coefficients but no control over the diffusion term. In this scenario, we first provide the well-posedness of the exploratory Hamilton–Jacobi–Bellman equation with linear growth coefficients and polynomial growth reward function. By such a well-posedess result we achieve PIA’s convergence by establishing a quantitative locally uniform <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="245_2025_10249_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="30" /> </InlineMediaObject> <EquationSource Format="TEX">\({{\mathcal {C}}}^{1,\alpha }\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mrow> <mi mathvariant="script">C</mi> </mrow> <mrow> <mn>1</mn> <mo>,</mo> <mi>α</mi> </mrow> </msup> </math></EquationSource> </InlineEquation> estimates for the generated value sequence.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Policy Iteration for Exploratory Hamilton–Jacobi–Bellman Equations

  • Hung Vinh Tran,
  • Zhenhua Wang,
  • Yuming Paul Zhang

摘要

We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where the controls applied to the diffusion term satisfy a smallness condition. We demonstrate the convergence of PIA based on a uniform \({{\mathcal {C}}}^{2,\alpha }\) C 2 , α estimate for the value sequence generated by PIA, and provide a quantitative convergence analysis for this scenario. Second, we investigate PIA with unbounded coefficients but no control over the diffusion term. In this scenario, we first provide the well-posedness of the exploratory Hamilton–Jacobi–Bellman equation with linear growth coefficients and polynomial growth reward function. By such a well-posedess result we achieve PIA’s convergence by establishing a quantitative locally uniform \({{\mathcal {C}}}^{1,\alpha }\) C 1 , α estimates for the generated value sequence.