Approach on Machine Learning Techniques for Anomaly-Based Web Intrusion Detection Systems: Using CICIDS2017 Dataset
摘要
Anomaly-based IDS (AIDS) is a vital component of network security, and its successful deployment relies on the utilization of deep learning or machine learning algorithms. However, a review of existing AIDS literature reveals various issues, such as arbitrary algorithm selection, parameter choices, testing criteria, utilization of outdated datasets, and inadequate result analysis and validation. This comprehensive paper addresses these shortcomings by benchmarking previous AIDS studies using CICIDS2017 dataset and attack types to identify the most suitable algorithms, parameters, and testing criteria. Aim is to assess two commonly used Machine Learning (ML) algorithms and present models for each of them, thoroughly examine training and tuning parameters, and to evaluate the performance of classifiers, it is essential to utilize various metrics including F-score T-positive, T-negative rates, accuracy, precision, and recall. To ensure an accurate evaluation, make use of the multiclass CICIDS2017 dataset, which is known for its high level of class imbalance. Comprising real-world network attacks for the training and testing time of ML-AIDS models to evaluate their efficiency. Findings demonstrate that Random Forest classifier and Decision Tree Classifier models outperform other approaches, exhibiting superior capabilities in detecting web attacks. Models surpass others in accuracy, precision, recall, and F-score, showcasing their remarkable effectiveness and efficiency. Specifically, the accuracy evaluation yielded a score of 98.13%, while the recall achieved an impressive 98.51%.