<p>Large language models (LLMs) have general capabilities that make them promising candidates for autonomous cyber defense (ACD). However, their reliability remains uncertain, and a thorough comparison with traditional reinforcement learning (RL) approaches is lacking. This paper presents a systematic evaluation of LLM-based defense agents in the CybORG++ environment. We compare multiple models and prompting strategies against a specific pretrained Proximal Policy Optimization (PPO) reinforcement learning baseline across multiple adversaries and network topologies. We find that the best-performing LLM agents can outperform the RL baseline over long horizons. They also avoid false positives when only benign traffic is present. However, LLM agents show higher variance across runs than the RL agent. In addition, their performance is sensitive to prompt design, and the best results relative to the RL baseline are obtained with prompts that combine examples with explicit, manually specified strategy rules. In contrast, with zero-shot prompts, LLM effectiveness drops. Zero-shot results can be improved by adding memory via autoregressive prompting and, for some models, by including explicit action feedback, but they do not match the best few-shot rule-based prompts. With regard to generalization, although LLMs can be applied to new adversaries or network topologies without major prompt changes, this transfer is not reliable. Finally, all LLM agents show higher latency than the RL policy. Overall, our results show that LLMs can produce competitive defense policies without fine-tuning but require manually engineered prompts. Furthermore, their higher variance and slower response times could render them unsuitable for some real-world scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A systematic evaluation of large language models for autonomous cyber defense

  • Thibaut Jacques

摘要

Large language models (LLMs) have general capabilities that make them promising candidates for autonomous cyber defense (ACD). However, their reliability remains uncertain, and a thorough comparison with traditional reinforcement learning (RL) approaches is lacking. This paper presents a systematic evaluation of LLM-based defense agents in the CybORG++ environment. We compare multiple models and prompting strategies against a specific pretrained Proximal Policy Optimization (PPO) reinforcement learning baseline across multiple adversaries and network topologies. We find that the best-performing LLM agents can outperform the RL baseline over long horizons. They also avoid false positives when only benign traffic is present. However, LLM agents show higher variance across runs than the RL agent. In addition, their performance is sensitive to prompt design, and the best results relative to the RL baseline are obtained with prompts that combine examples with explicit, manually specified strategy rules. In contrast, with zero-shot prompts, LLM effectiveness drops. Zero-shot results can be improved by adding memory via autoregressive prompting and, for some models, by including explicit action feedback, but they do not match the best few-shot rule-based prompts. With regard to generalization, although LLMs can be applied to new adversaries or network topologies without major prompt changes, this transfer is not reliable. Finally, all LLM agents show higher latency than the RL policy. Overall, our results show that LLMs can produce competitive defense policies without fine-tuning but require manually engineered prompts. Furthermore, their higher variance and slower response times could render them unsuitable for some real-world scenarios.