Machine learning-based detection of complex cyberattacks
摘要
With the increasing complexity of modern cyberattacks, such as advanced persistent threats, reconnaissance, and steganography, traditional rule-based and signature-based detection methods are becoming less effective. Machine learning (ML) provides advanced capabilities for identifying sophisticated and stealthy attacks by efficiently processing large volumes of data and uncovering hidden patterns. This paper presents a systematic review of existing approaches to complex cyberattack detection based on machine learning techniques, encompassing an analysis of 68 research articles. The review evaluates the performance of individual algorithms compared to ensemble approaches, examines commonly used ML methods, and analyzes datasets used in experimental studies. The results show that ensemble models generally outperform individual classifiers, with detection accuracy improvements ranging from 0.4 % to 28.52 %. Machine learning methods such as XGBoost, Random Forest, and LightGBM are identified as particularly effective across various attack types. Supervised learning remains dominant, though interest in unsupervised and semi-supervised methods is increasing to address novel threats. Frequently used datasets include NSL-KDD, UNSW-NB15, and newer APT-focused datasets such as DAPT2020 and SCVIC-APT-2021. The findings confirm the strong potential of ML for adaptive and proactive cybersecurity systems.