The Role of Hyperparameter Tuning in Phishing Website Classification: A Comparative Analysis of ML Models
摘要
Phishing website classification is crucial in cybersecurity due to the rapid growth of online platforms. This study employs various Machine Learning (ML) algorithms, including Light Gradient Boosting Machine (LGBM), Quadratic Discriminant Analysis (QDA), AdaBoost, Logistic Regression (LR) and Decision Tree (DT)., to classify phishing websites. Performance evaluation includes both hyperparameter-tuned and non-tuned models using the publicly available “Phishing Websites Dataset” from Mendeley, featuring 58,000 websites, including 37,647 phishing sites and 111 features. Hyperparameter-tuned ML algorithms achieve over 89% accuracy, with LGBM scoring the highest at 97%. Untuned ML algorithms achieve over 73% accuracy, with LGBM leading at 96%. Hyperparameter tuning boosts LGBM's accuracy by 1%, underscoring its importance in phishing website classification within cybersecurity.