错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PruVer: Verification Assisted Pruning for Deep Reinforcement Learning

  • Briti Gangopadhyay,
  • Pallab Dasgupta,
  • Soumyajit Dey

摘要

Active deployment of Deep Reinforcement Learning (DRL) based controllers on safety-critical embedded platforms require model compaction. Neural pruning has been extensively studied in the context of CNNs and computer vision, but such approaches do not guarantee the preservation of safety in the context of DRL. A pruned network converging to high reward may not adhere to safety requirements. This paper proposes a framework, PruVer, that performs iterative refinement on a pruned network with verification in the loop. This results in a compressed network that adheres to safety specifications with formal guarantees over small time horizons. We demonstrate our method in model-free RL environments, achieving 40–60% compaction, significant latency benefits (3 to 10 times), and bounded guarantees for safety properties.