Evaluating Large Language Models on Cybersecurity Knowledge and Skills: A Comparative Analysis
摘要
Can large language models serve as valuable tools in cybersecurity tasks? This study explores the breadth of knowledge that large language models (LLMs), including GPT-4, LLaMA 3.1, Mistral-large, and Gemini-1.5, possess in the field of cybersecurity and their potential to understand and solve critical cybersecurity operations. These operations include tasks such as incident response, threat identification, and the ability to address complex issues related to cybersecurity defense and analysis. To rigorously evaluate this potential, we propose a systematic evaluation of these models by defining a benchmark framework. Our evaluation focuses on two main categories: knowledge (theoretical understanding) and skills (practical application), covering key areas of cybersecurity. The responses of the models are assessed using tailored metrics for each type of question. Our findings reveal variability in performance between the different LLMs and cybersecurity domains, highlighting both their strengths and limitations. This work offers a foundation for future research and a standardized approach to the evaluation of LLMs in cybersecurity.