This study presents a comparative analysis of machine learning and deep learning using two methods of text data preprocessing for the task of email classification into spam and non-spam. The chosen text preprocessing methods were TF-IDF and Word Embedding. The aim of the experiment was to identify the most effective models for classification tasks and determine the best combination of models with the selected text processing methods. The work used seven deep learning models and seven varieties of machine learning models. The findings revealed that whilst the Word Embedding technique is more appropriate for deep learning models, TF-IDF couples better with machine learning models. Using TF-IDF, Random Forest had the best accuracy among the machine learning models, a score of 0.9787. With scores ranging from 0.9601 to 0.9484, practically all deep learning models showed high accuracy and performance using Word Embedding. Nevertheless, the hybrid CNN-LSTM model efficiently manages classification problems independent of the selected text processing technique.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Machine Learning and Deep Learning Models for Email Spam Classification Using TF-IDF and Word Embedding Techniques

  • Kamronbek Yusupov,
  • Md Rezanur Islam,
  • Ibrokhim Muminov,
  • Mahdi Sahlabadi,
  • Kangbin Yim

摘要

This study presents a comparative analysis of machine learning and deep learning using two methods of text data preprocessing for the task of email classification into spam and non-spam. The chosen text preprocessing methods were TF-IDF and Word Embedding. The aim of the experiment was to identify the most effective models for classification tasks and determine the best combination of models with the selected text processing methods. The work used seven deep learning models and seven varieties of machine learning models. The findings revealed that whilst the Word Embedding technique is more appropriate for deep learning models, TF-IDF couples better with machine learning models. Using TF-IDF, Random Forest had the best accuracy among the machine learning models, a score of 0.9787. With scores ranging from 0.9601 to 0.9484, practically all deep learning models showed high accuracy and performance using Word Embedding. Nevertheless, the hybrid CNN-LSTM model efficiently manages classification problems independent of the selected text processing technique.