Nonlinear scalarization in stochastic multi-objective MDPs
摘要
Many sequential decision problems of interest are best described in relation to multiple, possibly conflicting objectives. In recent years, there has been a growing interest in the application of reinforcement learning techniques to solve this type of problem. However, most existing multi-objective reinforcement learning (MORL) algorithms only apply in cases where the utility over the different criteria is linear. Unfortunately, in many real-world use cases, users’ preferences cannot be represented accurately by linear scalarization. Nonlinear scalarizations provide the required expressivity, but pose a number of issues when tackled by reinforcement learning, with so far little study and no definite solution. In particular, current methods struggle to deal with the most general MORL setting, where the environment is stochastic and the scalarization of the expected returns associated with each objective is nonlinear. In this paper, we present a model-free, value-based algorithm aimed at this type of problem: we prove that it converges to the optimal deterministic policy, confirm this result with experimental evidence, and discuss its limitations.