Reinforcement Learning for Partially Observable Models
摘要
Partially Observed Markov Decision Problems (POMDPs) offer a practically rich and relevant, and mathematically challenging class of models. Even in the most basic setup of finite state-action models, the analysis and computation of optimal solutions is complicated. The existence of optimal policies has in general been established via converting, or reducing, the original partially observed stochastic control problem to a fully observed Markov Decision Problem (MDP) with probability measure valued (belief) states, leading to a belief-MDP. However, computing an optimal policy for this fully observed model, and thus for the original POMDP, using classical methods (such as dynamic programming, policy iteration, linear programming) is not simple even if the original system has finite state and action spaces, since the state space of the fully observed (reduced) model is always uncountable.