Evaluation of the maturity of LLMs in the cybersecurity domain
摘要
The increasing sophistication of cybersecurity threats demands innovative solutions for network defense and security education. This study evaluates the potential of Large Language Models (LLMs) in automating cybersecurity tasks, specifically focusing on three critical areas: Honeypot generation, Malware Detection, and Capture The Flag (CTF) exercise creation. We introduce two novel evaluation frameworks: the Cybersecurity Language Understanding (CSLU) benchmark, which assesses model knowledge through domain-specific multiple-choice questions, and an automated evaluation system that measures models’ ability to generate functional security artifacts. Using these frameworks, we evaluated seven state-of-the-art LLMs, including GPT-4, Gemini Pro, and Claude 3 Opus. Results demonstrate that current LLMs exhibit strong capabilities in Malware analysis, with four models achieving perfect scores. However, performance varied significantly in CTF exercise generation, indicating areas for improvement. GPT-4, Gemini Pro, and Claude 3 Opus consistently outperformed other models across all tasks. Performance patterns suggest that model size correlates with task effectiveness, though architecture-specific optimizations also play a significant role. Our findings indicate that LLMs can effectively automate certain cybersecurity tasks, particularly in Malware Detection and analysis. However, their capabilities vary across different security domains, suggesting the need for specialized training or domain-specific adaptations for optimal performance.