错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Guarding Inboxes: An NLP-Based Approach for Email Spam Detection

  • Linda Varghese,
  • Rajesh R. Pai,
  • Nandini Kumari,
  • G. Savitha,
  • S. Girisha

摘要

The unsolicited and misleading email material is sent in bulk to many recipients, sometimes known as spam or junk email. In the recent time, the increasing volume of such emails poses challenges to the electronic communication. This study therefore tries to build a reliable and accurate method for identifying and preventing spam emails, for improving user experience and information security. The dataset “Spam email classification” extracted from the Kaggle website is used in this study to detect and categorize email spam. It analyzes the text of the email using natural language processing and applies machine learning techniques to original unbalanced and resampled balanced datasets. The results indicate that the random forest model performs most effectively with an F1-score of 98% and an accuracy of 93%, respectively.