This work introduces a flexible and adaptive approach to URL classification, aiming to develop a tool capable of learning from diverse datasets to accurately identify malicious URLs. The scientific novelty of this approach lies in developing an ensemble of criteria that, with greater accuracy than previously proposed methods, enables the formulation of an optimal criterion for detecting phishing messages. The main idea is to design a system that remains effective across datasets from different sources, ensuring its robustness in real-world cybersecurity applications. To achieve this, a broad selection of modern classification algorithms was evaluated, allowing the identification of the three most efficient models. These top-performing classifiers were then combined into a VotingClassifier, leveraging ensemble learning to enhance predictive accuracy, reduce variance, and improve overall model stability. The study follows a supervised classification approach, where models are trained on labelled URL data (e.g., “malicious” or “benign”) to classify previously unseen URLs. The application of ensemble methods addresses common challenges in URL classification, such as data imbalance, noisy features, and varying dataset structures. By integrating multiple classifiers, the system compensates for individual model weaknesses, leading to more reliable predictions. The experimental results confirm that the strategic use of data mining algorithms, combined with rigorous data preprocessing and feature engineering, provides a powerful and scalable solution for detecting malicious URLs, effectively adapting to evolving cybersecurity threats, and serving as a valuable asset in network security and threat detection systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a High-Precision Model for Detecting Malicious Domain Names in Anti-spam Systems Using Artificial Intelligence Technologies

  • Petro Venherskyi,
  • Volodymyr Lesyk

摘要

This work introduces a flexible and adaptive approach to URL classification, aiming to develop a tool capable of learning from diverse datasets to accurately identify malicious URLs. The scientific novelty of this approach lies in developing an ensemble of criteria that, with greater accuracy than previously proposed methods, enables the formulation of an optimal criterion for detecting phishing messages. The main idea is to design a system that remains effective across datasets from different sources, ensuring its robustness in real-world cybersecurity applications. To achieve this, a broad selection of modern classification algorithms was evaluated, allowing the identification of the three most efficient models. These top-performing classifiers were then combined into a VotingClassifier, leveraging ensemble learning to enhance predictive accuracy, reduce variance, and improve overall model stability. The study follows a supervised classification approach, where models are trained on labelled URL data (e.g., “malicious” or “benign”) to classify previously unseen URLs. The application of ensemble methods addresses common challenges in URL classification, such as data imbalance, noisy features, and varying dataset structures. By integrating multiple classifiers, the system compensates for individual model weaknesses, leading to more reliable predictions. The experimental results confirm that the strategic use of data mining algorithms, combined with rigorous data preprocessing and feature engineering, provides a powerful and scalable solution for detecting malicious URLs, effectively adapting to evolving cybersecurity threats, and serving as a valuable asset in network security and threat detection systems.