<p>This paper studies mean-field Markov decision processes (MDPs) with centralised stopping under non-exponential discount functions. The problem differs fundamentally from most existing studies on mean-field optimal control/stopping due to its time-inconsistency by nature. We look for subgame-perfect relaxed equilibria, namely randomised stopping policies that satisfy time-consistent planning with future selves from the perspective of a social planner. On the other hand, unlike many previous studies on time-inconsistent stopping where decreasing impatience plays a key role, we are interested in a general discount function without imposing any conditions. As a result, a study on relaxed equilibria becomes necessary as pure-strategy equilibria need not exist in general. We formulate relaxed equilibria as fixed points of a complicated operator; their existence by a direct method is challenging. To overcome the obstacles, we first introduce an auxiliary problem under an entropy regularisation on the randomised policy and discount function and establish the existence of regularised equilibria as fixed points to an auxiliary operator via the Schauder fixed-point theorem. Next, we show that the regularised equilibria converge as the regularisation parameter&#xa0;<InlineEquation ID="IEq1"> <EquationSource Format="MATHML"><math> <mi>λ</mi> </math></EquationSource> <EquationSource Format="TEX">$\lambda $</EquationSource> </InlineEquation> tends to&#xa0;0; the limit corresponds to a fixed point for the original operator and hence is a relaxed equilibrium. We also establish some connections between the mean-field MDP and the <InlineEquation ID="IEq2"> <EquationSource Format="MATHML"><math> <mi>N</mi> </math></EquationSource> <EquationSource Format="TEX">$N$</EquationSource> </InlineEquation>-agent MDP when&#xa0;<InlineEquation ID="IEq3"> <EquationSource Format="MATHML"><math> <mi>N</mi> </math></EquationSource> <EquationSource Format="TEX">$N$</EquationSource> </InlineEquation> is sufficiently large in our time-inconsistent&#xa0;setting.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Time-inconsistent mean-field stopping problems: a regularised equilibrium approach

  • Xiang Yu,
  • Fengyi Yuan

摘要

This paper studies mean-field Markov decision processes (MDPs) with centralised stopping under non-exponential discount functions. The problem differs fundamentally from most existing studies on mean-field optimal control/stopping due to its time-inconsistency by nature. We look for subgame-perfect relaxed equilibria, namely randomised stopping policies that satisfy time-consistent planning with future selves from the perspective of a social planner. On the other hand, unlike many previous studies on time-inconsistent stopping where decreasing impatience plays a key role, we are interested in a general discount function without imposing any conditions. As a result, a study on relaxed equilibria becomes necessary as pure-strategy equilibria need not exist in general. We formulate relaxed equilibria as fixed points of a complicated operator; their existence by a direct method is challenging. To overcome the obstacles, we first introduce an auxiliary problem under an entropy regularisation on the randomised policy and discount function and establish the existence of regularised equilibria as fixed points to an auxiliary operator via the Schauder fixed-point theorem. Next, we show that the regularised equilibria converge as the regularisation parameter  λ $\lambda $ tends to 0; the limit corresponds to a fixed point for the original operator and hence is a relaxed equilibrium. We also establish some connections between the mean-field MDP and the N $N$ -agent MDP when  N $N$ is sufficiently large in our time-inconsistent setting.