错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Partially Observable Markov Chains

  • Julio B. Clempner,
  • Alexander Poznyak

摘要

The controlled Partly Observable Markov Decision Process (POMDP) architecture has shown to be effective in a variety of fields where one is required to disclose only partial knowledge about the problem’s structure and parameters. Since some state variables are difficult to track and measure correctly, it may be more beneficial to base judgments on less accurate information. This chapter focuses on the design of an observer for a class of partially observable ergodic homogeneous finite Markov chains. The major objective of the suggested approach is to derive the formulas for computing an observer and, as a consequence, the best control strategy. We create a new variable that combines the policy, the observation kernel, and the distribution vector in order to solve the issue. To retrieve the important variables, we derive the formulas. The POMDP model’s parameters are being learned in a dynamic context in this work. The development of the adaptive policies is based on an identification method, in which we count the number of unobserved events to estimate the components of the utility and transition matrices. The practical applications of the theoretical concerns addressed to a portfolio optimization problem are demonstrated through the use of a numerical example.