Analysis of Machine Learning Approach for Spamming Electronic Mail Detection
摘要
Individuals and organizations use electronic mail (E-mail) to send and receive digital messages over the internet. In 2023, 347.3 billion emails are sent and received per day among which are spam emails. These emails can lead to communication overload, waste of time, irritation, loss of important emails, and potential exposure to malware. Irrelevant or unsolicited messages, malware, advertisements etc. are common features of spam emails. In this study we are employing various machine learning methods such as logistic regression, decision trees, random forest, SVM, naive bayes and BERT to identify spam emails. We transform email data into a specialized format known as a sparse matrix using techniques like TFIDF. Then, we apply these methods, encompassing logistic regression, decision trees, random forests, and naive Bayes, to aid us in the identification process. Additionally, the text is tokenized using BERT Transformer, serving as input for BERT. Considering multiple parameters, Logistic Regression and SVM Algorithms 99% accuracy was achieved in spam detection.