A Survey of Machine Learning Techniques in Phishing Detection
摘要
Phishing attacks, a prevalent method for illicitly acquiring individuals’ sensitive information from the internet, pose significant threats to users’ security. These attacks, orchestrated by hackers, involve the theft of protected passwords, private details, and even financial transactions, resulting in stolen money. Typically, perpetrators of phishing attacks manipulate and conceal well-known, legitimate websites to deceive users into divulging their personal data. To counteract such cyber threats, numerous websites and cybersecurity experts employ various techniques. Whitelisting and blacklisting, along with heuristic algorithms based on visual resemblance, constitute some of the prevalent strategies. However, this study proposes an advanced approach—a machine learning-based categorization technique enriched with heuristic features. These features are derived from critical characteristics such as the uniform resource locator (URL), source code, session details, security type employed, protocol in use, and the type of site being accessed. The proposed model utilizes five distinct machine learning techniques, including random forest, decision trees, and logistic regression, to comprehensively evaluate its efficacy. By leveraging these advanced methodologies, this study aims to enhance the accuracy and efficiency of phishing detection, ultimately fortifying defenses against these malicious online activities.