错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification of Phishing URLs Using Machine Learning Models

  • Meghashyam Vivek,
  • Nithin Premjith,
  • Aaron Antonio Johnson,
  • Ashutosh Kumar Maurya,
  • I. Diana Jeba Jingle

摘要

In this study, we provide a machine learning-based method for identifying phishing URLs. Sixteen features, including Have IP, Have At, URL Length, URL Depth, Non-standard double slash, HTTPS domain, Shortened URL, Hyphen Count, DNS Record, Domain age, Domain active, iFrame, Mouse Over, Right click, Web Forwards, and Label, were extracted from the 600,000 URLs we gathered as a dataset of legitimate and phishing URLs. We then used this dataset to train a variety of machine learning models. These included standalone models such Naive Bayes, Logistic Regression, Decision Trees, and K-Nearest Neighbors (KNN). We also used ensemble models like Hard Voting, XGBoost, Random Forests, and AdaBoost. Finally, we used deep learning models such as Artificial Neural Networks (ANN), Long Short-Term Memory (LSTM), Gated Recurrent Units (GRU) and Convolutional Neural Networks (CNN). On evaluation of performance metrics like accuracy, precision, recall, train time and prediction time it was found that XGBoost provides the best performance across all categories.