错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Generation of Optimal Attack Strategies Based on Reinforcement Learning in Complex Network Environments

  • Shoujia Chen,
  • Jianyu Deng

摘要

The efficient generation of attack strategies is a critical research problem in the field of cybersecurity offense and defense. Well-designed attack strategies can significantly enhance penetration efficiency. As network environments become increasingly complex, strategy generation must take a broader range of factors into account, making this task even more challenging. Deep reinforcement learning (DRL) has demonstrated outstanding performance in the field of automated decision-making and is increasingly being applied to address practical decision-making problems in cybersecurity. However, the lack of realistic and complex network training environments hinders research on the generation of attack strategies. Additionally, the sparse reward distribution in such environments poses challenges to the effective training of agents. The simplistic design of training environments in previous studies has led to attack agents generating strategies that are difficult to generalize to new environments. Therefore, enabling attack agents to learn and explore in multi-feature network environments and applying their strategies to new environments is a more practical approach. This paper designs an agent training environment that incorporates multi-feature characteristics closely aligned with real-world network scenarios and proposes an attack strategy planning framework based on DRL. Additionally, we designed a dynamic reward mechanism based on the characteristics of the environment to solve the problem of low training efficiency caused by traditional sparse reward distribution. Finally, in the experiments, we validated the rationality of the proposed environment. After a certain number of training iterations, the agent was able to generate effective attack strategies. We also conducted comparative experiments with various common DRL algorithms and thoroughly analyzed the performance of each model based on changes in cumulative rewards.