Multi-agent Simulation and Reinforcement Learning to Optimize Moving Target Defense
摘要
As cyber threats continue to evolve, Moving Target Defense (MTD) strategies have emerged as a promising approach to enhancing network security by dynamically altering system configurations. However, optimizing MTD requires balancing security improvements with system availability. In this work, we propose a framework for optimizing MTD strategies, leveraging reinforcement learning (RL) and a cybersecurity simulation environment (named CybORG-MTD), to train defensive agents. Our approach introduces a reward function that explicitly models the trade-off between security and availability, enabling RL agents to learn effective defense policies. Through empirical evaluations, we demonstrate that our proposed methodology outperforms existing techniques in optimizing MTD strategies. While our results highlight the effectiveness of RL-based cyber defense, we also discuss key challenges, including scalability, and adaptive attacker behaviors.