Feature-Based Explainable Reinforcement Learning in Environments with Multiple Sources of Risk
摘要
Explainable Reinforcement Learning is key in bringing the current neural network-based state-of-the-art reinforcement learning methods to real-world environments. In particular, explaining the risks of the agent’s decision-making process is critical to deploying such models in safety-critical tasks. The previous feature-based methods for characterizing risk in reinforcement learning settings were not well suited for handling situations when there are multiple sources of risk. This work attempts to address this shortcoming by providing a post-hoc method that can explain multiple sources of risk. Our experiments show that the proposed method can provide more insights into the workings of the agent while avoiding the issues faced by the previous work in multi-risk environments.