Phishing attacks continue to pose significant risks in cybersecurity, requiring advanced detection mechanisms to combat evolving threats effectively. This paper presents a comprehensive approach for phishing URL detection, integrating ensemble learning, hyperparameter tuning, and Explainable AI techniques. Leveraging a dataset from Kaggle comprising 100,078 instances and 20 features, various machine learning classifiers were trained and fine-tuned using hyperparameter optimization methods. Among these models, AdaBoost, XGBoost, and Random Forest emerged as the top three performers. These three models were subsequently combined to form a tri-model ensemble, capitalizing on their complementary strengths to enhance accuracy and robustness. Integration of Explainable AI techniques, such as LIME and SHAP, enhances transparency and interpretability in the ensemble model’s predictions. Results demonstrate significant performance gains, with the ensemble model achieving an accuracy of 92.8%. The interpretive analysis offers insights into prediction factors, fostering trust, and informed decision-making in cybersecurity. This framework holds promise for bolstering defenses against phishing attacks and advancing cybersecurity efforts in an increasingly complex threat landscape.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Tri-model Ensemble Approach for Phishing URL Detection Using Explainable AI and Hyperparameter Tuning Techniques

  • M. Bharath,
  • S. Gowtham,
  • Ashwini Kodipalli,
  • Trupthi Rao

摘要

Phishing attacks continue to pose significant risks in cybersecurity, requiring advanced detection mechanisms to combat evolving threats effectively. This paper presents a comprehensive approach for phishing URL detection, integrating ensemble learning, hyperparameter tuning, and Explainable AI techniques. Leveraging a dataset from Kaggle comprising 100,078 instances and 20 features, various machine learning classifiers were trained and fine-tuned using hyperparameter optimization methods. Among these models, AdaBoost, XGBoost, and Random Forest emerged as the top three performers. These three models were subsequently combined to form a tri-model ensemble, capitalizing on their complementary strengths to enhance accuracy and robustness. Integration of Explainable AI techniques, such as LIME and SHAP, enhances transparency and interpretability in the ensemble model’s predictions. Results demonstrate significant performance gains, with the ensemble model achieving an accuracy of 92.8%. The interpretive analysis offers insights into prediction factors, fostering trust, and informed decision-making in cybersecurity. This framework holds promise for bolstering defenses against phishing attacks and advancing cybersecurity efforts in an increasingly complex threat landscape.