As cyber threats continue to evolve, Moving Target Defense (MTD) strategies have emerged as a promising approach to enhancing network security by dynamically altering system configurations. However, optimizing MTD requires balancing security improvements with system availability. In this work, we propose a framework for optimizing MTD strategies, leveraging reinforcement learning (RL) and a cybersecurity simulation environment (named CybORG-MTD), to train defensive agents. Our approach introduces a reward function that explicitly models the trade-off between security and availability, enabling RL agents to learn effective defense policies. Through empirical evaluations, we demonstrate that our proposed methodology outperforms existing techniques in optimizing MTD strategies. While our results highlight the effectiveness of RL-based cyber defense, we also discuss key challenges, including scalability, and adaptive attacker behaviors.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-agent Simulation and Reinforcement Learning to Optimize Moving Target Defense

  • William Valentine,
  • Etienne Borde,
  • Mengmeng Ge

摘要

As cyber threats continue to evolve, Moving Target Defense (MTD) strategies have emerged as a promising approach to enhancing network security by dynamically altering system configurations. However, optimizing MTD requires balancing security improvements with system availability. In this work, we propose a framework for optimizing MTD strategies, leveraging reinforcement learning (RL) and a cybersecurity simulation environment (named CybORG-MTD), to train defensive agents. Our approach introduces a reward function that explicitly models the trade-off between security and availability, enabling RL agents to learn effective defense policies. Through empirical evaluations, we demonstrate that our proposed methodology outperforms existing techniques in optimizing MTD strategies. While our results highlight the effectiveness of RL-based cyber defense, we also discuss key challenges, including scalability, and adaptive attacker behaviors.