The Dual-Edged Sword of Large Language Models in Phishing
摘要
Background: With the rise of Large Language Model (LLM) technologies and LLM-based chatbots like ChatGPT, Copilot or Gemini, cyberattacks such as phishing are getting more sophisticated by using AI to craft personalized phishing messages. This poses a challenge for cybersecurity. Aim: This study explores the complexities of AI-enhanced phishing strategies, their success factors, and how LLMs can be used to improve cybersecurity defenses against phishing. Method: We delve into how LLMs, especially GPT 3.5 and 4, can detect and combat phishing. By experimenting with prompting techniques such as zero-shot, multi-shot, and chain-of-thought, we assess how these models fare in spotting phishing emails across various datasets. Results: The findings show that while GPT-4 demonstrates high precision and recall, the decision to deploy LLMs must consider cost-effectiveness, given their computational demand and operational costs.