A Discount Vanishing Approximation for Markov Decision Processes with Risk Sensitivity
摘要
In this paper optimal control of risk-sensitive Markov decision processes with countable states is studied. The state space is not assumed to be communicating. The focus is on dependence of the optimal values on the transition characteristics-communication, transience or absorption. A vanishing discount approach is used to introduce a partition of the state space, and certain transformation of the optimal values under discount is shown to convergence to the optimal values under risk sensitivity, as the discount factor tends to vanish. The partition of the state space turns out to be closely related to the characteristics of state communication, but weights more on the values under discount.