The vast amount of malicious traffic on the network in a critical Information infrastructure environment requires scrutiny and a proper understanding of APT features and structure before deployment in developing predictive machine learning models (MLM). Thorough comprehension is beneficial to avoid erroneous and biased model development due to improper datasets fed into the models. This chapter explores the performance of various conventional MLM and ensemble models in predicting APT attacks with an imbalanced APT dataset. The study uses the SCVIC 2021 APT dataset. An exploratory experiment was undertaken to understand the importance of carefully performing preprocessing, mostly data balancing, for a multi-class problem. Results demonstrate the performance of K-neighbors (KNN), random forest (RF), logistic regression, support vector machine (SVM), extra tree classifier, and XGBoost on imbalanced multi-class data posed by the selected APT dataset. Accuracy, balanced accuracy score, precision, recall, F1-score, and geometric mean score were evaluated. Results elucidate that feeding imbalanced data into state-of-the-art models produces biased results, resulting in the majority class always being favored. Results for accuracy and balanced accuracy demonstrate a greater difference for all models; the accuracy rate is 1.00 due to the data imbalance, while balanced accuracy for XGBoost and extra tree classifier is 0.88 with the geometric mean values of 0.91. Support vector machine is the worst model for both balanced accuracy and geometric mean with 0.17 and 0.00, respectively. Precision, recall, and F1-measure for all the model predictions for class 3, which is “normal traffic,” are 1.00, which needs further investigation as that is the majority class, implying some bias in the results. The chapter confirms as proof of concept that accuracy is not a good measure for imbalanced multi-class problems and that the applicability of balanced accuracy and geometric means presents a meaningful performance evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-class Classification on Network Traffic Data: Evaluation with a Multi-class APT Dataset

  • Mamoqenelo P. Morolong,
  • Fungai Bhunu Shava,
  • Attlee M. Gamundani

摘要

The vast amount of malicious traffic on the network in a critical Information infrastructure environment requires scrutiny and a proper understanding of APT features and structure before deployment in developing predictive machine learning models (MLM). Thorough comprehension is beneficial to avoid erroneous and biased model development due to improper datasets fed into the models. This chapter explores the performance of various conventional MLM and ensemble models in predicting APT attacks with an imbalanced APT dataset. The study uses the SCVIC 2021 APT dataset. An exploratory experiment was undertaken to understand the importance of carefully performing preprocessing, mostly data balancing, for a multi-class problem. Results demonstrate the performance of K-neighbors (KNN), random forest (RF), logistic regression, support vector machine (SVM), extra tree classifier, and XGBoost on imbalanced multi-class data posed by the selected APT dataset. Accuracy, balanced accuracy score, precision, recall, F1-score, and geometric mean score were evaluated. Results elucidate that feeding imbalanced data into state-of-the-art models produces biased results, resulting in the majority class always being favored. Results for accuracy and balanced accuracy demonstrate a greater difference for all models; the accuracy rate is 1.00 due to the data imbalance, while balanced accuracy for XGBoost and extra tree classifier is 0.88 with the geometric mean values of 0.91. Support vector machine is the worst model for both balanced accuracy and geometric mean with 0.17 and 0.00, respectively. Precision, recall, and F1-measure for all the model predictions for class 3, which is “normal traffic,” are 1.00, which needs further investigation as that is the majority class, implying some bias in the results. The chapter confirms as proof of concept that accuracy is not a good measure for imbalanced multi-class problems and that the applicability of balanced accuracy and geometric means presents a meaningful performance evaluation.