Evolution of Malicious URL Detection: A Review of Techniques for Malicious URL Detection and Classification
摘要
The ever-increasing number of cyber threats also means that malicious URLs are one such attack vector that has emerged as a vector for malware dissemination, phishing, and data theft. Blacklisting and rule-based filtering, traditional detection methods, have been rendered less effective by the constantly changing nature of these threats. To combat this, researchers have created sophisticated detection and classification methods using machine learning (ML), deep learning (DL), and natural language processing (NLP). This work provides an exhaustive survey on advancing malicious URL detection methods and discusses how malicious URL detection techniques have evolved from heuristic and signature-based methods to state-of-the-art artificial intelligence (AI)-based frameworks. We review a set of methodologies, from feature-based ML models to deep neural networks and hybrid techniques that combine several detection frameworks. Furthermore, we address some of the major challenges including adversarial evasion strategies, scalability, and limitations on real-time detection. They also review how dataset curation, feature selection, and evaluation metrics can contribute to detection accuracy. This review synthesizing the recently made advances provides insights into various approaches, including strengths and limitations, and paves the way ahead for future research directions. Our focus is to provide data on adaptive, scalable, and robust models that could be made use of to address any emerging threats in the cyber world. This study helps in developing more efficient and resilient malicious URL detection systems.