This research examines the application of natural language processing (NLP) techniques to improve email classification, focusing on effectively identifying spam and ham (legitimate) emails. Given the growing risks spam emails present to cybersecurity and communication, developing reliable classification methods is critical. The study incorporates both modern and traditional techniques: FastText, recognized for its speed and precision in text classification, and the classic Bag-of-Words model paired with a random forest classifier. Preprocessing steps such as tokenization, stopword removal, and stemming were utilized to enhance feature extraction and model performance. Fast-Text, with its ability to manage large datasets and understand word relationships, was tested on the datasets, where it demonstrated excellent accuracy in differentiating between spam and non-spam emails. Combining Fast-text with traditional models offers a comprehensive approach to spam detection. The results indicate that these methods can significantly increase the resilience of email filtering systems, providing an efficient and scalable approach to combating the ever-evolving threat of spam. This research contributes valuable insights into improving email security by integrating both traditional and modern NLP strategies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Natural Language Processing in Email Spam Filtering: A Comparative Study

  • Jay Parmar,
  • Dev Parikh,
  • Amit Thakkar,
  • Dhaval Bhoi

摘要

This research examines the application of natural language processing (NLP) techniques to improve email classification, focusing on effectively identifying spam and ham (legitimate) emails. Given the growing risks spam emails present to cybersecurity and communication, developing reliable classification methods is critical. The study incorporates both modern and traditional techniques: FastText, recognized for its speed and precision in text classification, and the classic Bag-of-Words model paired with a random forest classifier. Preprocessing steps such as tokenization, stopword removal, and stemming were utilized to enhance feature extraction and model performance. Fast-Text, with its ability to manage large datasets and understand word relationships, was tested on the datasets, where it demonstrated excellent accuracy in differentiating between spam and non-spam emails. Combining Fast-text with traditional models offers a comprehensive approach to spam detection. The results indicate that these methods can significantly increase the resilience of email filtering systems, providing an efficient and scalable approach to combating the ever-evolving threat of spam. This research contributes valuable insights into improving email security by integrating both traditional and modern NLP strategies.