Background: With the rise of Large Language Model (LLM) technologies and LLM-based chatbots like ChatGPT, Copilot or Gemini, cyberattacks such as phishing are getting more sophisticated by using AI to craft personalized phishing messages. This poses a challenge for cybersecurity. Aim: This study explores the complexities of AI-enhanced phishing strategies, their success factors, and how LLMs can be used to improve cybersecurity defenses against phishing. Method: We delve into how LLMs, especially GPT 3.5 and 4, can detect and combat phishing. By experimenting with prompting techniques such as zero-shot, multi-shot, and chain-of-thought, we assess how these models fare in spotting phishing emails across various datasets. Results: The findings show that while GPT-4 demonstrates high precision and recall, the decision to deploy LLMs must consider cost-effectiveness, given their computational demand and operational costs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Dual-Edged Sword of Large Language Models in Phishing

  • Alec Siemerink,
  • Slinger Jansen,
  • Katsiaryna Labunets

摘要

Background: With the rise of Large Language Model (LLM) technologies and LLM-based chatbots like ChatGPT, Copilot or Gemini, cyberattacks such as phishing are getting more sophisticated by using AI to craft personalized phishing messages. This poses a challenge for cybersecurity. Aim: This study explores the complexities of AI-enhanced phishing strategies, their success factors, and how LLMs can be used to improve cybersecurity defenses against phishing. Method: We delve into how LLMs, especially GPT 3.5 and 4, can detect and combat phishing. By experimenting with prompting techniques such as zero-shot, multi-shot, and chain-of-thought, we assess how these models fare in spotting phishing emails across various datasets. Results: The findings show that while GPT-4 demonstrates high precision and recall, the decision to deploy LLMs must consider cost-effectiveness, given their computational demand and operational costs.