To synthesize a safe and optimal controller for switched hybrid systems, one can first synthesize a shield that ensures safety, and then apply reinforcement learning within the constraints of the shield to obtain the desired controller. However, developing such a shield for switched hybrid systems typically requires a full model of the environment, which is not always available. Instead, historical data of the environment might be available. In this paper, we introduce a method for the construction of safety shields based on different scenarios captured in historical data. We show how individual shields for different scenarios can be combined to obtain a single shield that is provably safe within the bounds of the observed scenarios. We demonstrate the method using an industrial case study of a stormwater detention pond, which includes ten years of historical data of different rain events/scenarios. Our experimental results show that the shielded optimal controller ensures safety across all individual historical rain scenarios compared to the unshielded optimal controller. Additionally, we empirically show that the shield may also generalize for scenarios not covered by the historical data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data-Driven Shielding of Online Reinforcement Learning: A Stormwater Pond Case Study

  • Esther Hahyeon Kim,
  • Martijn Angelo Goorden,
  • Kim Guldstrand Larsen,
  • Thomas Dyhre Nielsen

摘要

To synthesize a safe and optimal controller for switched hybrid systems, one can first synthesize a shield that ensures safety, and then apply reinforcement learning within the constraints of the shield to obtain the desired controller. However, developing such a shield for switched hybrid systems typically requires a full model of the environment, which is not always available. Instead, historical data of the environment might be available. In this paper, we introduce a method for the construction of safety shields based on different scenarios captured in historical data. We show how individual shields for different scenarios can be combined to obtain a single shield that is provably safe within the bounds of the observed scenarios. We demonstrate the method using an industrial case study of a stormwater detention pond, which includes ten years of historical data of different rain events/scenarios. Our experimental results show that the shielded optimal controller ensures safety across all individual historical rain scenarios compared to the unshielded optimal controller. Additionally, we empirically show that the shield may also generalize for scenarios not covered by the historical data.