In recent years, websites are collecting information from users for many purposes. But users are not ensuring that the collected information is used only for the anticipated purpose. Securing the information from unauthorized users and minimizing the risk of misusing the sensitive information is one way to ensure the security of data. Phishing is the common way to deceive the sensitive information like bank credentials or personal information from users and utilize the deceived information for malicious activities. This causes many hazardous effects on individuals as well as organizations. Securing the information from the phishers requires technological efforts and valuable for global cause. The proposed work mainly focuses on collecting the website URL through web crawler and validates the authenticity of the URL by calculating similarity score. Also, the proposed work employs on chi-square feature selection method for minimizing the number of features for classification purpose, which results in better response time with 0.12 s. Different classifiers are tested for evaluating proposed method. As per simulation and performance analysis, random forest classifier outperforms other conventional methods by achieving accuracy of 95.36% without similarity score index and enhanced accuracy of 96.07% after including similarity score index. Additionally, detection time is computed to analyze the time complexity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Phishing URL Classification Based on Machine Learning with Keyword Specific Web Crawler

  • Hallimysore Devaraj Nandeesha,
  • Bantiganalli Thimappa Prasanna

摘要

In recent years, websites are collecting information from users for many purposes. But users are not ensuring that the collected information is used only for the anticipated purpose. Securing the information from unauthorized users and minimizing the risk of misusing the sensitive information is one way to ensure the security of data. Phishing is the common way to deceive the sensitive information like bank credentials or personal information from users and utilize the deceived information for malicious activities. This causes many hazardous effects on individuals as well as organizations. Securing the information from the phishers requires technological efforts and valuable for global cause. The proposed work mainly focuses on collecting the website URL through web crawler and validates the authenticity of the URL by calculating similarity score. Also, the proposed work employs on chi-square feature selection method for minimizing the number of features for classification purpose, which results in better response time with 0.12 s. Different classifiers are tested for evaluating proposed method. As per simulation and performance analysis, random forest classifier outperforms other conventional methods by achieving accuracy of 95.36% without similarity score index and enhanced accuracy of 96.07% after including similarity score index. Additionally, detection time is computed to analyze the time complexity.