With the growing complexity of cloud environments, static firewall policies often fall short in addressing the dynamic nature of modern cyber threats. This study explores the application of reinforcement learning (RL) in developing adaptive firewall policies that can autonomously adjust to evolving network conditions and threats. By employing a multi-objective reward shaping approach, this research aims to optimise firewall performance across conflicting objectives, such as maximising security, minimising false positives, and maintaining network performance. Using a deep Q-learning framework, the RL-based firewall dynamically adapts its policies in response to live traffic, adjusting its actions based on a reward function tailored to multi-objective optimisation. Simulated experiments demonstrate that the proposed system effectively balances security and performance, showing higher attack detection rates with minimal impact on legitimate traffic flow. The study’s findings underscore the potential of multi-objective reward shaping in creating robust, adaptive firewall policies and pave the way for future exploration in scalable, self-learning security solutions for cloud platforms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Firewall Policies for Cloud Security: A Multi-objective Reinforcement Learning Approach

  • Aptin Babaei,
  • Abbas Khosravi,
  • Ibrahim Hossain

摘要

With the growing complexity of cloud environments, static firewall policies often fall short in addressing the dynamic nature of modern cyber threats. This study explores the application of reinforcement learning (RL) in developing adaptive firewall policies that can autonomously adjust to evolving network conditions and threats. By employing a multi-objective reward shaping approach, this research aims to optimise firewall performance across conflicting objectives, such as maximising security, minimising false positives, and maintaining network performance. Using a deep Q-learning framework, the RL-based firewall dynamically adapts its policies in response to live traffic, adjusting its actions based on a reward function tailored to multi-objective optimisation. Simulated experiments demonstrate that the proposed system effectively balances security and performance, showing higher attack detection rates with minimal impact on legitimate traffic flow. The study’s findings underscore the potential of multi-objective reward shaping in creating robust, adaptive firewall policies and pave the way for future exploration in scalable, self-learning security solutions for cloud platforms.