Identification of Malicious URLs: A Purely Lexical Approach
摘要
Internet users are increasingly exposed to security vulnerabilities stemming from malicious Uniform Resource Locators (URLs), which act as conduits for cyber threats. These threats, often orchestrated by sophisticated cybercriminals, underscore the importance of comprehending the intricate dynamics involved to devise robust defense mechanisms. This scholarly exposition delineates an efficacious approach for discerning diverse categories of malicious URLs leveraging machine learning algorithms. Notably, our methodology obviates the necessity of directly accessing such URLs for extracting pertinent information, relying solely on attributes inherent within the lexical composition of the URLs. The empirical analyses are predicated on meticulously curated datasets from reputable repositories such as Kaggle and PhishTank, culminating in competitive performance vis-à-vis existing literature that predominantly focuses on network-centric or content-based features.