Spam Detection and Classification Based on Ensemble Methods Using Natural Language Processing
摘要
SPAM refers to any form of undesired and unsolicited electronic communication that is distributed in mass. Spam may be communicated in many ways. Email is the most common medium for the transmission of spam, although it can also be transmitted via text message, phone call, or through social media. In our research work, we discuss email spam. It is undesired digital communication sent to an individual, group, or corporation. These Spam emails might threaten the email user. By collecting online signup addresses spammers reveal the Attacks on authentic users. The attackers create fake accounts to spam and cause the issues like password theft, storage issues, phishing, and more cyber frauds. Therefore, there is a need to detect and filter these spam emails. In this paper, To detect email spam we use Natural Language-processing over the ensemble machine learning methods like Bagging, GradientBoosting, and ADABoosting. Our results show good accuracy of 88.99%, 96.57%, and 95.78% by the ensemble machine learning algorithms such as Bagging, GradientBoosting, and ADABoosting respectively.