Can large language models serve as valuable tools in cybersecurity tasks? This study explores the breadth of knowledge that large language models (LLMs), including GPT-4, LLaMA 3.1, Mistral-large, and Gemini-1.5, possess in the field of cybersecurity and their potential to understand and solve critical cybersecurity operations. These operations include tasks such as incident response, threat identification, and the ability to address complex issues related to cybersecurity defense and analysis. To rigorously evaluate this potential, we propose a systematic evaluation of these models by defining a benchmark framework. Our evaluation focuses on two main categories: knowledge (theoretical understanding) and skills (practical application), covering key areas of cybersecurity. The responses of the models are assessed using tailored metrics for each type of question. Our findings reveal variability in performance between the different LLMs and cybersecurity domains, highlighting both their strengths and limitations. This work offers a foundation for future research and a standardized approach to the evaluation of LLMs in cybersecurity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Large Language Models on Cybersecurity Knowledge and Skills: A Comparative Analysis

  • Mohamed El Amine Bekhouche,
  • Myria Bouhaddi,
  • Wahib Larbi,
  • Kamel Adi

摘要

Can large language models serve as valuable tools in cybersecurity tasks? This study explores the breadth of knowledge that large language models (LLMs), including GPT-4, LLaMA 3.1, Mistral-large, and Gemini-1.5, possess in the field of cybersecurity and their potential to understand and solve critical cybersecurity operations. These operations include tasks such as incident response, threat identification, and the ability to address complex issues related to cybersecurity defense and analysis. To rigorously evaluate this potential, we propose a systematic evaluation of these models by defining a benchmark framework. Our evaluation focuses on two main categories: knowledge (theoretical understanding) and skills (practical application), covering key areas of cybersecurity. The responses of the models are assessed using tailored metrics for each type of question. Our findings reveal variability in performance between the different LLMs and cybersecurity domains, highlighting both their strengths and limitations. This work offers a foundation for future research and a standardized approach to the evaluation of LLMs in cybersecurity.