The exponential rise of malicious activities on the internet underscored the critical need for robust detection mechanisms to safeguard users from potential threats. In this paper, the authors propose an innovative method for enhancing malicious URL detection by utilizing ensemble fusion techniques that integrate both ML and DL methodologies. The proposed method began by loading and preprocessing a large-scale dataset comprising 5,49,346 URLs sourced from Kaggle. Through feature engineering and extraction, the dataset is transformed into a numerical format suitable for model training, employing TF-IDF to capture the importance of features. Subsequently, individual ML models are trained, including Random Forest, XGBoost, and Gradient Boosting, as well as the DL models Multi-Layer Perceptron (MLP), RNN, LSTM, and GRU, on the preprocessed data. Random Forest achieved a recall of 97% and an accuracy of 97.50%, while LSTM demonstrated a recall and accuracy of 97% and 97.50%, respectively. Then, ensemble fusion techniques, specifically stacking and the meta-learner approach, were used to combine the predictions from all individual models and produce a final prediction. Through comprehensive evaluation and performance analysis, the proposed method demonstrated the efficacy of ensemble fusion model in accurately detecting malicious URLs, achieving superior performance compared to individual models. The proposed ensemble model with logistic regression as a meta-learner achieved an accuracy of 98.4% and a recall of 98%. These findings underscore the robustness and superior performance of the ensemble fusion approach in accurately identifying malicious URLs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble Fusion for Enhanced Malicious URL Detection by Integrating Machine Learning and Deep Learning Techniques

  • Raja Rao PBV,
  • Kiran Sree Pokkuluri,
  • M. Prasad,
  • Neeraj Sharma,
  • BSatya Narayana Murthy,
  • Adina Karunasri

摘要

The exponential rise of malicious activities on the internet underscored the critical need for robust detection mechanisms to safeguard users from potential threats. In this paper, the authors propose an innovative method for enhancing malicious URL detection by utilizing ensemble fusion techniques that integrate both ML and DL methodologies. The proposed method began by loading and preprocessing a large-scale dataset comprising 5,49,346 URLs sourced from Kaggle. Through feature engineering and extraction, the dataset is transformed into a numerical format suitable for model training, employing TF-IDF to capture the importance of features. Subsequently, individual ML models are trained, including Random Forest, XGBoost, and Gradient Boosting, as well as the DL models Multi-Layer Perceptron (MLP), RNN, LSTM, and GRU, on the preprocessed data. Random Forest achieved a recall of 97% and an accuracy of 97.50%, while LSTM demonstrated a recall and accuracy of 97% and 97.50%, respectively. Then, ensemble fusion techniques, specifically stacking and the meta-learner approach, were used to combine the predictions from all individual models and produce a final prediction. Through comprehensive evaluation and performance analysis, the proposed method demonstrated the efficacy of ensemble fusion model in accurately detecting malicious URLs, achieving superior performance compared to individual models. The proposed ensemble model with logistic regression as a meta-learner achieved an accuracy of 98.4% and a recall of 98%. These findings underscore the robustness and superior performance of the ensemble fusion approach in accurately identifying malicious URLs.