错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Statistical Reinforcement Learning and Dynamic Treatment Regimes

  • Tao Shen,
  • Yifan Cui

摘要

This chapter introduces to readers the concept and methodology of reinforcement learning and modern-day dynamic treatment regimes in statistics. This discussion should be of interest to those who wish to go into the depth of statistical reinforcement learning. We start with introducing the Markov decision process. Then, several methods such as policy iteration, value iteration, temporal difference learning, and policy gradient are presented. Next, dynamic treatment regimes in both point exposure and time-varying settings are introduced. We also include a discussion of causal reinforcement learning which is an emerging field of policy learning. Several illustrative examples are provided for practitioners based on simulated data. This review aims to provide an overview of various state-of-the-art methods for reinforcement learning and dynamic treatment regimes. The sample codes provided in this chapter are available at https://github.com/taoshen2022/Statistical-Reinforcement-Learning-and-Dynamic-Treatment-Regimes-Supplementary-Material .