Identifying Phishing Attacks Using URL-Centric Approaches with Character-Conscious Transformer-Based Language Model
摘要
Phishing causes annual losses in the billions of dollars, seriously endangering the online economy. Email is the most common means of carrying out phishing attempts. Although the current research trends in phishing email detection have been covered in several review articles, it is important to tackle this problem from many perspectives. Except for one survey that briefly mentioned employing Natural Language Processing (NLP) algorithms for categorization while considering a few alternatives, none of these surveys have exhaustively examined the use cases of NLP techniques for spotting phishing. The objective of this paper is to diminish the existing limitations by meticulously reviewing kinds of literature and using NLP to the identification of phishing attacks. Using NLP and Universal Resource Locator (URL) parameters, the authors of this work suggested a method to identify Phishing-as-a-Service (PhaaS) attacks, which have been a recent problem. The experiment showed that the output includes a report summarizing how the model's performance was judged, together with figures for accuracy, precision, recall, and F1-score in the form of percentages 93%, 92%, 93%, and 93%, respectively, which is higher than what has previously been suggested. These results significantly outstrip previous benchmarks, signifying a substantial advancement in phishing detection technology.