With the advancement of technology, cyber threats have been on the rise, and with it, detection and avoidance of links at both personal and organizational levels are very important. One of the most dangerous threats is called phishing, where a link takes the user to a page meant to perform any form of harmful actions like stealing data, hacking into a system, or committing financial fraud. These attacks may result in business interruptions, theft of important information, and consequently lead to tremendous losses. The paper focuses on the use of machine learning in identifying harmful links. To further enhance the accuracy in detection, two powerful techniques of machine learning are used here: Random Forest (RF) and Gradient Boosting Machine (GBM). The first approach creates many decision trees at training, whereas the second builds the models one step at a time, correcting mistakes with each iteration. Both approaches find threats well. Training of such methods starts with two entirely different data sets. Every dataset contains features from varying sets of links. Training of the RF model will be done using one of these data sets, and the model GBM will be tested over the same data set. In this way, learning occurs in the models where they pick up on different sorts of patterns and become powerful in real situations. The study results indicate that RF and GBM models effectively identify harmful links and reduce the chances of false labeling of safe links as harmful and vice versa. This advancement is significant for quickly spotting possible cyber threats and reducing risks at the earliest. The paper also proposes a method of handling shortened links, where the actual destination link is not shown and may lead to malicious websites. The features that can be added for the improved detection of harmful content within these links, by hidden shortened links, include specifics related to shortened links.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Customized Fusion of RF and XGB for Enhanced Malicious URL Detection

  • A. Pramod Kumar,
  • P. Madhusudan,
  • R. Vasanth Naik,
  • S. Pranathi,
  • V. Abhijeet Kumar

摘要

With the advancement of technology, cyber threats have been on the rise, and with it, detection and avoidance of links at both personal and organizational levels are very important. One of the most dangerous threats is called phishing, where a link takes the user to a page meant to perform any form of harmful actions like stealing data, hacking into a system, or committing financial fraud. These attacks may result in business interruptions, theft of important information, and consequently lead to tremendous losses. The paper focuses on the use of machine learning in identifying harmful links. To further enhance the accuracy in detection, two powerful techniques of machine learning are used here: Random Forest (RF) and Gradient Boosting Machine (GBM). The first approach creates many decision trees at training, whereas the second builds the models one step at a time, correcting mistakes with each iteration. Both approaches find threats well. Training of such methods starts with two entirely different data sets. Every dataset contains features from varying sets of links. Training of the RF model will be done using one of these data sets, and the model GBM will be tested over the same data set. In this way, learning occurs in the models where they pick up on different sorts of patterns and become powerful in real situations. The study results indicate that RF and GBM models effectively identify harmful links and reduce the chances of false labeling of safe links as harmful and vice versa. This advancement is significant for quickly spotting possible cyber threats and reducing risks at the earliest. The paper also proposes a method of handling shortened links, where the actual destination link is not shown and may lead to malicious websites. The features that can be added for the improved detection of harmful content within these links, by hidden shortened links, include specifics related to shortened links.