错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing the Performance Variations of Naive Bayes, Linear SVM, and Random Forest for Spam Detection: A Comprehensive Study on the &Quot; Spam or Ham" Dataset

  • Bhawna Ojha,
  • Pradeep Yadav,
  • Rakhi Arora,
  • Nitin Dixit,
  • Gaurav Dubey,
  • Khemchand Shakyawar

摘要

Spam detection is the process of identifying and filtering unsolicited or unwanted messages, often referred to as spam, from legitimate or desired messages. It is an important task in email systems, messaging platforms, and other communication channels to protect users from unwanted or potentially harmful content. The performance of Naive Bayes, Linear SVM, and Random Forest classification machine learning algorithms on the “Spam or Ham” dataset was evaluated using various metrics, including accuracy, precision, recall, and F1 score. The results indicate that all three classifiers are suitable for this dataset, but Naive Bayes outperforms Random Forest and Linear SVM in terms of precision, recall, and F1 score. This suggests that Naive Bayes can accurately classify spam messages while minimizing the number of false positives. The optimal classifier will depend on features such as dataset size, the specific business problem, and available computing resources. However, based on this specific experiment’s results, Naive Bayes appears to be a good choice for classifying spam messages. The accuracy of 0.98 is achieved by Naive Bayes algorithm.