Effective Email Filtering: A Machine Learning Approach to Distinguish Safe and Phishing Emails with SVM
摘要
Email communication is an essential part of our daily routines, but it also opens the door to cyber threats, with phishing attacks being a significant concern. This research focuses on a crucial task: categorizing emails into two groups—those deemed “Safe” and those flagged as “Phishing.” This study focuses on the development and evaluation of machine learning models designed for precise email classification. The study utilized a curated dataset sourced from Kaggle, encompassing 18,650 diverse email samples. Emphasizing evaluation robustness, the reported accuracy, precision, recall, and F1-score metrics were meticulously derived from testing on a separate validation set, ensuring the reliability and generalizability of our findings. Various techniques, such as logistic regression, naive Bayes, support vector machines (SVMs), k-nearest neighbors (KNNs), decision trees (DT), random forests (RFs), and K-means clustering, were employed in our comprehensive analysis. The primary metrics considered during performance evaluation include accuracy, precision, recall, and F1-score. Among the models tested, SVM emerged as the standout performer, achieving an impressive accuracy rate of 96.91% in differentiating between safe and phishing emails. Notably, SVM demonstrated robust precision and recall, positioning it as a leading candidate for effective email classification. Logistic regression and random forests also exhibited strong performances, with accuracies of 96.83% and 96.35%, respectively. Even the unsupervised method of K-means clustering showed promise in email categorization. This research significantly contributes to the field of email security by highlighting the effectiveness of diverse machine learning models in classifying emails. It underscores the critical need for employing robust techniques to combat phishing attacks and ensure the security of email communication. The findings of this study lay a foundation for enhancing email security measures. They provide valuable insights for developing more precise email filters, ultimately safeguarding users from phishing threats. The practical implications of this research are substantial, benefiting both organizations and individuals in their endeavors to maintain secure email communications. The results underscore the importance of continued research and innovation in the field of machine learning for bolstering email security.